Setting Up a Chi-Square Test for Homogeneity or Independence
A two-way table can answer two different kinds of questions. Learn to choose the test from the study design, write clear hypotheses and check whether a chi-square model is appropriate.
By the end of this lesson, you should be able to:
- Distinguish comparing distributions across populations from testing association within population.
- Choose homogeneity or independence using how the data were collected.
- Write hypotheses that name the categorical variable or variables and the population or populations.
- Check randomization, independence of observations, the condition when needed and every expected cell count.
- Describe chi-square curves and how their shape changes with degrees of freedom.
- Explain what a complete test setup establishes and what still needs a calculation.
Before you start: Review two-way tables, conditional proportions and Topic 3.13’s population-test reasoning. Each individual should contribute count to a clearly defined row-and-column combination.
First time learning this? Follow the three-school preference study, then compare it with a single-school survey.
Here to revise? Use the method-and-conditions checklist, then attempt the practice questions before opening the solutions.
The concept in 60 seconds
Researchers independently select random samples from Schools A, B and C. Each student chooses preferred notes format: digital only, printed only or a mixed format. The researchers ask whether the distribution of preferences is the same across the school populations.
This is a chi-square test for homogeneity. “Homogeneity” means sameness: under the null, the populations share the same preference distribution. The observed samples need not have identical counts or percentages.
Independent samples at Schools A, B and C; notes-format preference in each. Choose homogeneity.
academy sample; grade level and notes-format preference recorded for every student. Choose independence.
Write the appropriate population null and alternative, then justify the design and every expected cell.
A two-way table organizes the counts. Its dimensions alone do not tell you whether the populations were separately sampled or whether sample was classified by variables.
Now imagine a different design: random sample from a single school, with each student classified by grade level and preferred notes format. The question is whether those categorical variables are associated in that school’s population. This is a chi-square test for independence.
The key choice: comparing categorical distribution across separately sampled populations or assigned treatments → homogeneity. Looking for an association between categorical variables in sampled population → independence. Start with the design and question.
All numerical studies on this page are fictional teaching examples. This lesson builds and justifies a test plan; Topic 3.15 uses that plan to calculate the test statistic, -value and conclusion.
Quick check: does a table automatically mean a test for independence?
No. The same table dimensions can come from independent samples of populations or sample classified by variables. The sampling design and research question determine the test name.
Choose the test from the study design
Homogeneity: compare population or treatment distributions
Take independent random samples from or more populations and measure the same categorical response in each. Alternatively, randomly assign experimental units to treatments and compare the categorical outcome distributions. Name the populations or treatments and the response variable.
In the school study, samples of students are taken separately from A, B and C. The variable is preferred notes format, with the same categories at every school.
| Population | Digital only | Printed only | Mixed | Total |
|---|---|---|---|---|
| School A | ||||
| School B | ||||
| School C | ||||
| Total |
Different sample sizes make raw-count comparisons misleading. For example, digital preferences out of at A is ; out of at B is . A homogeneity null concerns the population category proportions, rather than equality of sample counts.
Independence: relate variables in population
Take random sample from population, then record categorical variables for each individual. For example, sample students at academy and record grade level (9, 10 or 11) and preferred notes format.
Samples are separately taken from the schools. Compare the distribution of preferred notes format across populations: homogeneity.
sample from an academy’s grades 9–11 population is classified by grade and format. Ask about association: independence.
Both designs can produce the same cell counts and expected-count arithmetic. Their questions, hypotheses and sampling checks still refer to different populations and designs.
| Design and question | Method | Name in your setup |
|---|---|---|
| Independent random samples from several populations; compare the same categorical response distributions. | Chi-square test for homogeneity | The categorical response and all sampled populations. |
| Random assignment to treatments; compare categorical outcome distributions. | Chi-square test for homogeneity | The categorical outcome and assigned treatments. |
| random sample; investigate association between categorical variables. | Chi-square test for independence | Both categorical variables and the single target population. |
| The same people before and after; investigate a change in marginal proportions. | Requires an appropriate paired-change method | The linked observations and the change question; do not invent independent groups. |
| Quantitative measurements such as test-score means. | Use an appropriate quantitative-data method | The population quantitative outcome; categorical counts answer a different question. |
Keep the counting unit clear
A cell count is the number of individuals in a row-and-column combination. Give each individual exactly category for each variable. If students select several separate options, counting each selection as a different student changes the data structure. A “mixed format” category can work when it is clearly defined, exclusive response.
sample can contain several groups after classification. Finding grade-level groups inside a single SRS does not turn it into separately collected samples. Conversely, two-way row labels alone do not reveal how observations were selected.
Quick check: SRS of city residents records age group and preferred transport mode. Which test fits?
A chi-square test for independence, if its conditions hold. The categorical variables are age group and preferred transport mode, and the target is their association in the city-resident population.
Write contextual null and alternative hypotheses
Write hypotheses about the population distributions or population association. Ordinary wording is often clearer than forcing a long list of symbols.
| Test | Null | Alternative |
|---|---|---|
| Homogeneity | The categorical response distribution is the same across the populations or treatments. | The distributions are not all the same; at least differs. |
| Independence | The categorical variables are independent in the target population. | The categorical variables are associated in that population. |
Homogeneity: the three-school study
: The distribution of preferred notes format—digital only, printed only or mixed—is the same among all students at Schools A, B and C.
: The distribution of preferred notes format is not the same across all school populations; at least population distribution differs.
The null means that the digital proportion is the same across schools, the printed proportion is the same across schools and the mixed proportion is the same across schools. It does not require the categories to be equally likely. A shared distribution could be digital, printed and mixed.
The alternative does not claim that every school differs from every other school or that every category differs. At least difference somewhere is enough for the general alternative.
Independence: a single-school survey
: Grade level and preferred notes format are independent, or have no association, among all students in grades 9–11 at the sampled academy.
: Grade level and preferred notes format are associated among those students.
Under independence, the population distribution of notes format is the same across grade levels. Under association, knowing the grade level changes the distribution of format preference.
Keep meanings of independence separate. Independent observations are a design requirement. Independence of the categorical variables is the null claim being investigated. Assuming the variables are independent as a condition would assume the answer to the research question.
Quick check: does the homogeneity alternative mean that all population distributions must be different?
No. It means they are not all the same. population differing from another in at least category is enough; the setup does not identify which category differs.
Understand expected cell counts
An observed count is what the sample actually contains. An expected count is the count predicted by the null model using the table’s marginal totals.
To check the expected-count condition, calculate each interior cell using:
This formula applies to both homogeneity and independence. It combines the null-model category proportions with the relevant group size, while preserving the row and column totals.
Under the homogeneity null, estimate common category share from all samples.
Use this row’s actual size. Unequal school samples have unequal expected counts.
Observed is . The expected count describes the null model, rather than an equal-category assumption.
Equivalent arithmetic: . The combined category shares are digital, printed and mixed; homogeneity allows these shares to be unequal.
School A’s digital-format cell
The A row total is ; the digital column total is ; the grand total is . Therefore:
The observed count is . Expected means that A’s sample of would have digital preferences under the estimated common digital rate. The difference of is one piece of sample evidence; it is not yet a test decision.
| Population | Digital only | Printed only | Mixed | Total |
|---|---|---|---|---|
| School A | ||||
| School B | ||||
| School C | ||||
| Total |
The interior expected counts are ; ; and . The smallest is . Expected counts can be decimals; do not round them to whole people before checking a threshold or calculating a statistic.
Two useful arithmetic checks
- Expected counts in each row add to the original row total.
- Expected counts in each column add to the original column total; all cells together add to the grand total.
Do not put a row percentage into the table as if it were a count. If only percentages are given, you need the relevant sample sizes to obtain the count data.
Quick check: why is A’s expected digital count , rather than ?
The null assumes the same distribution across populations, not equal shares across categories. The combined data estimate the digital share as , so A’s expected digital count is .
Check randomization, independence and expected counts
1. Randomization and independent observations
For homogeneity, identify the independent random samples or the appropriate randomized experiment. For independence, identify the random sample from the single population. Describe the actual design given in the problem.
Observations must be appropriately independent. The same person should not be treated as several independent people, and linked groups require suitable analysis. A huge voluntary website poll does not become an SRS because it has many responses. Count checks do not repair recruitment bias.
2. The condition when sampling without replacement
For homogeneity with separately sampled finite populations, check for each population individually. Schools A, B and C contain students:
- School A: .
- School B: .
- School C: .
For independence with sample, check the whole sample against the target population. For the academy sample of from students, check . Row groups found after sampling do not require invented population sizes.
The sampling condition is not required merely for random assignment. A separate finite-population sampling stage can still require its own check. Do not compare each assigned group with of the participant pool. If a sampling problem omits population sizes, state what would be needed to justify the check; do not invent a size.
3. Every expected cell count must exceed
For the 2026–27 AP Statistics curriculum used here, every interior expected count should be greater than . The official CED, Topic 3.14.D.1.iii (printed page 111), states this condition. Check the expected cells, excluding marginal totals.
Our school table passes because all expected counts exceed ; the minimum is . A total sample of alone is not the justification.
Boundary reminder: follow in this lesson. An exact expected count of does not meet that strict wording. Some other texts use ; keep the criterion consistent with the curriculum or task you are following. A value below does not become adequate by rounding.
Only the smaller sample’s rare-category cell is plotted. Row totals are and , column totals and , and grand total . The expected-count table is ; . Independent SRSs and both checks pass in this example; the expected-count condition does not.
In the separate small-group example, the expected counts are . The first fails. All observed counts are at least , but that is not the required check.
An observed does not automatically cause failure when its expected count exceeds and that cell is genuinely possible. A structurally impossible category combination is a different model issue; do not treat it as an ordinary chance observation.
Quick check: expected counts exceed and is . Is the condition met?
No. Every interior expected cell count must exceed under the criterion used here. State the failed cell instead of claiming that a large grand total makes the condition acceptable.
Understand chi-square curves
The symbol is read “chi-square.” A chi-square statistic measures how far observed counts are from expected counts relative to the expected counts.
Preview:
The sum includes every interior cell.
Each squared difference contributes a nonnegative amount when . Thus the statistic cannot be negative. It can equal when every observed count exactly matches its expected count. Larger values indicate a greater overall discrepancy from the null model.