Inference for Categorical Data: Proportions
Move from sample percentages to justified statements about populations. Learn how to estimate proportions with confidence intervals, test claims with p-values and analyze categorical distributions with chi-square tests.
Before you start: Review random sampling and experimental design from Unit 1, plus conditional probability, binomial models and sampling distributions from Unit 2.

Explore all 15 topics
Choose a topic to open its study notes, worked examples, visuals and practice. Follow the cards in order, or jump to the skill you want to review.
Understand estimates and sampling
Connect the population proportion, sample proportion and repeated-sampling distribution.
- Topic 3.1
Estimators
Use a statistic to estimate a population parameter. Distinguish bias from sampling variability and judge an estimator over repeated samples.
Study Topic 3.1 - Topic 3.2
Sampling Distributions for Sample Proportions
Describe the sampling distribution of a sample proportion. Find its center and spread, then justify any normal approximation.
Study Topic 3.2
Estimate one population proportion
Build an interval, interpret confidence and assess a proposed value.
- Topic 3.3
Constructing a Confidence Interval for a Population Proportion
Construct a one-proportion confidence interval. Check conditions, calculate uncertainty and identify the population proportion being estimated.
Study Topic 3.3 - Topic 3.4
Justifying a Claim Based on a Confidence Interval for a Population Proportion
Interpret the interval and confidence level correctly. Use the endpoints to assess a proposed population proportion.
Study Topic 3.4
Test a claim about one proportion
Set up hypotheses, understand p-values, calculate a test and consider possible errors.
- Topic 3.5
Setting Up a Test for a Population Proportion
State contextual hypotheses about one population proportion. Choose the test and check conditions using the null proportion.
Study Topic 3.5 - Topic 3.6
p-Values
Interpret a p-value assuming the null hypothesis is true. Match extremeness and tail direction to the alternative hypothesis.
Study Topic 3.6 - Topic 3.7
Carrying Out a Test for a Population Proportion
Carry out a one-proportion z test. Report the statistic and p-value, make a decision and answer the question in context.
Study Topic 3.7 - Topic 3.8
Potential Errors When Performing Tests
Describe Type I and Type II errors in context. Connect error probabilities, significance level and power to the decision.
Study Topic 3.8
Compare two population proportions
Model a sample difference, estimate a population difference or test a comparison.
- Topic 3.9
Sampling Distributions for the Difference Between Sample Proportions
Describe the sampling distribution of a difference in sample proportions. Preserve group order and combine variability correctly.
Study Topic 3.9 - Topic 3.10
Constructing a Confidence Interval for the Difference Between Two Population Proportions
Construct an interval for the difference between two population proportions. Check both samples and use the appropriate standard error.
Study Topic 3.10 - Topic 3.11
Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Proportions
Interpret the sign and endpoints of a difference interval. Assess a claim while distinguishing percentage points from percent change.
Study Topic 3.11 - Topic 3.12
Setting Up a Test for the Difference Between Two Population Proportions
Set up a two-proportion test. State the population comparison and justify the pooled null model and conditions.
Study Topic 3.12 - Topic 3.13
Carrying Out a Test for the Difference Between Two Population Proportions
Calculate the pooled two-proportion z statistic and p-value. Draw a conclusion that respects the design and group order.
Study Topic 3.13
Analyze categorical count tables
Select the chi-square design, verify conditions and complete the inference.
- Topic 3.14
Setting Up a Chi-Square Test for Homogeneity or Independence
Choose chi-square homogeneity or independence from the design. State contextual hypotheses and check every expected count.
Study Topic 3.14 - Topic 3.15
Carrying Out a Chi-Square Test for Homogeneity or Independence
Calculate expected counts, cell contributions, degrees of freedom and the right-tail p-value. Write a qualified population conclusion.
Study Topic 3.15
The big picture
A school wants to know the proportion of all its students who prefer digital notes. A random sample gives an estimate, but another sample could give a different percentage. An interval describes estimation uncertainty; a test evaluates a specific population claim. Comparing several schools or formats leads to a different categorical test.
Choose the method from the question, parameter and study design before calculating. A calculator output becomes a statistical argument only when you justify the conditions and explain the result for the population of interest.
By the end of this unit, you should be able to:
- Distinguish a population proportion from the sample proportion used to estimate it.
- Describe sampling variability and justify the appropriate approximation.
- Construct and interpret one-proportion and two-proportion confidence intervals.
- State hypotheses, interpret p-values and make contextual test decisions.
- Explain Type I and Type II errors and the meaning of power.
- Select and carry out chi-square tests for homogeneity or independence.
Your study roadmap
Use these stages as a learning order. If you are revising, choose the stage that matches the skill you need to practise. Each stage links back to its topic cards above.
- Stage 01 · Topics 3.1–3.2
Understand estimates and sampling
Connect the population proportion, sample proportion and repeated-sampling distribution.
- Stage 02 · Topics 3.3–3.4
Estimate one population proportion
Build an interval, interpret confidence and assess a proposed value.
- Stage 03 · Topics 3.5–3.8
Test a claim about one proportion
Set up hypotheses, understand p-values, calculate a test and consider possible errors.
- Stage 04 · Topics 3.9–3.13
Compare two population proportions
Model a sample difference, estimate a population difference or test a comparison.
- Stage 05 · Topics 3.14–3.15
Analyze categorical count tables
Select the chi-square design, verify conditions and complete the inference.
For each lesson: read the key idea, work through an example, try practice without the solution, then compare your reasoning with the explanation. Revisit the step you missed before moving on.
Revise with a purpose
Choose the route that fits your goal. While revising, practise explaining the method and result; a correct number on its own may leave the question unanswered.
Learning for the first time
Follow all five stages in order. Pair setup lessons with calculation and interpretation lessons: a test is incomplete without conditions and a contextual conclusion. Open the starting lesson →
Revising estimation
Use Topics 3.3–3.4 for one proportion or Topics 3.10–3.11 for a difference. Name the parameter, check conditions and explain the interval endpoints and confidence level separately. Open the starting lesson →
Revising tests
Use Topics 3.5–3.8 for one proportion, 3.12–3.13 for two proportions, or 3.14–3.15 for count tables. Start by identifying the design and writing the null and alternative in context. Open the starting lesson →
Choose the inference method
One proportion: estimate
If the question asks for a plausible range for one population’s success proportion p, use a one-proportion z interval when justified. The estimate and standard error use the observed sample proportion.
One proportion: test
If the question evaluates a specified proportion p₀ in one population, use a one-proportion z test when justified. The null proportion supplies the reference model and standard deviation for the test.
Two proportions: estimate
If the question asks for a plausible difference p₁ − p₂, use a two-proportion z interval when justified. Preserve group order, and use separate sample estimates in the interval standard error.
Two proportions: test
For a test of equal population proportions, use the pooled estimate under H₀. This test calculation differs from the unpooled interval calculation; choose the standard error to match the task.
Several populations or treatments
A chi-square test for homogeneity compares the distribution of one categorical response across independently sampled populations or appropriate randomly assigned treatment groups.
One population, two categorical variables
A chi-square test for independence evaluates association between two categorical variables recorded on the same sampled individuals. Independent observations are a design condition, not the null claim about the variables.
Common mistakes to catch
- Describing a p-value as the probability that H₀ is true
A p-value is calculated assuming H₀ is true. Describe a statistic at least as extreme in the direction specified by Hₐ. Review Topic 3.6
- Using the same standard error for every proportion calculation
Intervals use sample-based estimates. A one-proportion test uses p₀; a test of two equal proportions uses the pooled null estimate. Review Topic 3.12
- Checking observed counts for a chi-square approximation
Check each expected interior cell. The linked chi-square lessons use the current framework’s strict requirement E > 5. Review Topic 3.14
- Concluding that a large p-value proves equality or independence
Failing to reject means the data do not provide sufficient evidence against H₀ at the chosen significance level. Review Topic 3.15
Check your understanding
These are original, short retrieval questions with fictional teaching situations. Try each one before opening the answer. If your explanation is incomplete, follow the review link for the relevant topic.
Question 1
A random sample of 400 school students contains 160 who prefer digital notes. Identify p̂ and explain how it differs from p.
Check answer 1
The sample proportion is p̂ = 160/400 = 0.40. It describes this sample. The parameter p is the unknown proportion of all students in the target school population who prefer digital notes. Samples vary, so the estimate need not equal the parameter.
Question 2
A valid 95% confidence interval for the school’s digital-notes proportion is (0.42, 0.50). State its margin of error and interpret the interval.
Check answer 2
The midpoint is 0.46 and the margin of error is 0.04, or four percentage points. We are 95% confident that the interval from 0.42 to 0.50 captures the population proportion of the school’s students who prefer digital notes. The confidence level concerns the long-run capture rate of the interval procedure, not 95% of individual students.
Question 3
A justified two-sided test of H₀: p = 0.50 versus Hₐ: p ≠ 0.50 gives p-value 0.0047. At α = 0.05, what is the decision and what does the p-value describe?
Check answer 3
Since 0.0047 < 0.05, reject H₀. The data provide evidence that the population proportion differs from 0.50. Assuming p = 0.50, the p-value is the probability of a test statistic at least as extreme as the observed statistic in either direction under the justified null model. It is not the probability that p equals 0.50.
Question 4
A test fails to reject H₀: p = 0.50, but the true population proportion is actually 0.60. Which error occurred?
Check answer 4
This is a Type II error: failing to reject a false null hypothesis. In context, the test failed to detect that the population proportion differs from 0.50. A Type I error would instead reject H₀ when the null claim is true.
Question 5
A valid 95% interval for p₁ − p₂ is (−0.08, 0.02). What does this interval say about a possible difference of zero?
Check answer 5
Zero is a plausible value because it is inside the interval. The interval does not establish a nonzero difference. Plausible differences range from group 1 being eight percentage points lower to two percentage points higher than group 2. This does not prove that the population proportions are equal.
Question 6
Independent random samples from three schools produce a 3 × 3 format-preference table. Conditions are satisfied, with all expected cells greater than 5. χ² = 7.6875 and p = 0.10372. Name the test, find df and conclude at α = 0.05.
Check answer 6
Use a chi-square test for homogeneity. df = (3 − 1)(3 − 1) = 4. Since 0.10372 > 0.05, fail to reject H₀. The samples do not provide sufficient evidence that preferred-notes-format distributions differ among the three school populations. This is not proof that the distributions are identical.
Are you ready to move on?
Use this as a checklist: explain each item aloud or on paper without looking at the notes. A checked box is a reminder for your study session, not an assessment score.
If an item is not yet comfortable, choose the matching stage in the roadmap and retry that lesson’s practice. If these explanations are clear, work on mixed questions where the topic is not named for you.
Continue learning
Keep building your statistics skills
Unit 4 continues inference with quantitative data. Its opening topic is Sampling Distributions for Sample Means.
The unit sequence follows the AP Statistics course framework effective Fall 2026. Use this page to navigate NUM8ERS lessons and plan your revision.