AP Statistics / Unit 3: Inference for Categorical Data: Proportions / Topic 3.10
NUM8ERS study notes · Topic 3.10

Constructing a Confidence Interval for the Difference Between Two Population Proportions

22 groups can have different success rates. Use their sample results to estimate the population difference, then build an interval that shows the uncertainty in that estimate.

2026–27 curriculum6 worked examples10 practice questions6 visual guides

By the end of this lesson, you should be able to:

  • Choose a two-sample zz-interval and define both population proportions in context.
  • Check the random-sampling design, independent groups, both 10%10\% conditions and all 44 observed counts.
  • Calculate the sample difference and its unpooled standard error.
  • Find a critical zz-value, margin of error and confidence-interval endpoints.
  • Explain the difference in percentage points and keep the group order consistent.
  • Describe how confidence level and sample sizes affect interval width.

Before you start: Review Topic 3.9’s sampling distribution for a difference and Topic 3.3’s one-proportion interval. You should know p^=xn\hat{p}=\frac{x}{n} and the idea of estimate±margin of error\text{estimate}\pm\text{margin of error}.

First time learning this? Follow the school example through the conditions, standard error and five construction steps.

Here to revise? Use the interval checklist, then try the questions before opening hints or solutions.

The concept in 60 seconds

Imagine comparing the proportion of students who prefer digital study notes at 22 schools. Surveying every student may be impractical, so take a random sample from each school.

The sample proportions give 11 estimate of the difference. Because another pair of samples could give another estimate, report a confidence interval for the population difference rather than treating the observed difference as exact.

Point estimate±margin of error\text{Point estimate}\pm\text{margin of error}

(p^1−p^2)±z∗SE⁡(\hat p_1-\hat p_2)\pm z^*\operatorname{SE}
  • p^1−p^2\hat{p}_{1}-\hat{p}_{2}: the observed sample-proportion difference.
  • SE⁡\operatorname{SE}: an estimate of how much sample differences vary.
  • z∗z^*: the multiplier for the chosen confidence level.
  • z∗×SE⁡z^* \times \operatorname{SE}: the margin of error on either side of the estimate.

The interval estimates p1−p2p_{1}-p_{2}. The population proportions are unknown. The sample proportions are known after collecting the data and are used to construct the interval.

Visual guide 1: 22 samples estimate 11 population difference
Group 1 · School Ap^1=124200=0.62\hat{p}_{1}=\frac{124}{200}=0.62

This observed rate estimates the unknown population proportion p1p_{1} who prefer digital notes.

Group 2 · School Bp^2=69150=0.46\hat{p}_{2}=\frac{69}{150}=0.46

This observed rate estimates the unknown population proportion p2p_{2} with the same preference.

Keep A minus BPoint estimate: 0.160.16

Build an interval for p1−p2p_{1}-p_{2} around this sample difference, allowing for uncertainty from both groups.

A success means “prefers digital notes” in both samples. The school results are fictional. The 22 population proportions are unknown; the sample proportions are observed estimates.

Quick check: what is the interval trying to estimate?

The difference between the 22 population success proportions, p1−p2p_{1}-p_{2}, in a stated order. It is not estimating the already observed sample difference.

Define the 22 populations

All examples on this page are fictional teaching scenarios. For the main example, 22 independent SRSs are selected without replacement: 200200 students from School A’s 10,00010{,}000 students and 150150 students from School B’s 8,0008{,}000 students. Both samples answer the same yes/no question about preferring digital notes.

Observed results: 124124 students in the School A sample and 6969 in the School B sample prefer digital notes. The true population proportions are unknown.

Population parameters and sample statistics for the school example.
SymbolMeaningSchool example
p1p_{1}Proportion of all School A students who prefer digital notesUnknown population proportion
p2p_{2}Proportion of all School B students who prefer digital notesUnknown population proportion
p1−p2p_{1}-p_{2}Target population difference, A minus BUnknown parameter to estimate
x1x_{1}, n1n_{1}School A success count and sample size124124 and 200200
x2x_{2}, n2n_{2}School B success count and sample size6969 and 150150
p^1\hat{p}_{1}, p^2\hat{p}_{2}Separate sample proportions0.620.62 and 0.460.46
p^1−p^2\hat{p}_{1}-\hat{p}_{2}Point estimate of the population difference0.160.16, or 1616 percentage points
p^1=124200=0.62\hat p_1=\frac{124}{200}=0.62
p^2=69150=0.46\hat p_2=\frac{69}{150}=0.46
p^1−p^2=0.62−0.46=0.16\hat p_1-\hat p_2=0.62-0.46=0.16

The observed preference proportion is 1616 percentage points higher at School A. This sample difference estimates the unknown population difference; it does not establish that the true difference is exactly 0.160.16.

Choose the procedure before using a formula

Use a two-sample zz-interval for a difference between population proportions when you want to estimate the difference in success rates for 22 appropriately independent groups, with the conditions below satisfied.

  • The response is categorical, coded as success or failure.
  • There are 22 groups, each with a success count and sample size.
  • The goal is an interval estimate of p1−p2p_{1}-p_{2}.

If the response is a numerical measurement such as travel time, a proportion interval is not the right procedure. If the same people are measured 22 times, the data are paired and the independent two-sample formula does not handle that dependence.

Keep one subtraction order

Let group 1 be School A and group 2 be School B throughout. A positive A-minus-B difference means A has the higher proportion; a negative difference means B has the higher proportion.

Quick check: is p1p_{1} “the proportion in the sample from School A”?

No. p1p_{1} is the proportion among all students at School A who prefer digital notes. The sample proportion is p^1=0.62\hat{p}_{1}=0.62. Always identify the population and the success outcome when defining a parameter.

The standard error of the difference

Topic 3.9 used stated population proportions to calculate a theoretical sampling-distribution SD. Here, the population proportions are unknown, so we substitute each group’s own sample proportion to estimate that SD.

SE⁡=p^1(1−p^1)n1+p^2(1−p^2)n2\operatorname{SE}=\sqrt{\frac{\hat p_1(1-\hat p_1)}{n_1}+\frac{\hat p_2(1-\hat p_2)}{n_2}}
Visual guide 2: use 22 separate estimates inside 11 SE⁡\operatorname{SE}
A contribution0.0011780.001178

0.62(1−0.62)200\frac{0.62(1 – 0.62)}{200} estimates the variance contribution from School A.

B contribution0.0016560.001656

0.46(1−0.46)150\frac{0.46(1 – 0.46)}{150} estimates the variance contribution from School B.

Add, then square rootSE⁡≈0.0532353\operatorname{SE}\approx 0.0532353

0.001178+0.001656\sqrt{0.001178+0.001656}. Match each sample proportion to its own sample size.

The 22 estimated variance terms add under the justified independent-group model. Their square root is the unpooled standard error in proportion-difference units, about 5.325.32 percentage points.

School A contribution:

0.62(0.38)200=0.001178\frac{0.62(0.38)}{200}=0.001178

School B contribution:

0.46(0.54)150=0.001656\frac{0.46(0.54)}{150}=0.001656
SE⁡=0.001178+0.001656=0.002834≈0.0532353\operatorname{SE}=\sqrt{0.001178+0.001656}=\sqrt{0.002834}\approx0.0532353

This SE⁡\operatorname{SE} estimates a typical sampling fluctuation of about 0.05320.0532, or 5.325.32 percentage points, in the sample-proportion difference around the true population difference. It is a measure of estimated spread, not the margin of error by itself.

Why do the 22 terms add?

Variation from either sample can change the difference. For independent groups, their variance contributions add, even though the proportions subtract. Take the square root only after adding those contributions.

This interval uses an unpooled SE⁡\operatorname{SE}. Use 0.620.62 for School A and 0.460.46 for School B. Do not combine the counts into 11 pooled proportion and substitute it into both terms.

Pooling can be relevant to a test that assumes equal population proportions. A confidence interval here estimates the difference without assuming equality. Subtracting 22 separate one-proportion intervals is also not the construction rule for this interval.

The usual AP formula omits finite-population corrections when samples are selected without replacement. The separate 10%10\% checks justify the usual approximation.

Quick check: should you subtract the 22 standard deviations?

No. Add the 22 estimated variance contributions inside 11 square root. Subtracting SDs could make uncertainty cancel incorrectly.

Check the conditions

1. Randomization and independent groups

For a sampling study, use 22 independent random samples. The school example explicitly specifies independent SRSs. A voluntary poll does not become a random sample merely by having many responses.

For an appropriate two-group experiment, use random assignment of treatments to experimental units. Check that a suitable independent-group design is described. Observations from the same people before and after treatment are paired and need a method that accounts for that dependence.

2. Check 10%10\% separately when sampling without replacement

n1≤0.10N1n_1\le0.10N_1 and n2≤0.10N2n_2\le0.10N_2.

School A: 200≤0.10(10,000)=1,000200\le0.10(10{,}000)=1{,}000.

School B: 150≤0.10(8,000)=800150\le0.10(8{,}000)=800.

Both pass. This supports treating responses within each sample as approximately independent for the usual formula. A combined sampling fraction across both populations cannot replace these 22 checks.

The finite-population sampling check is not required merely because participants are randomly assigned to 22 experimental groups. Random sampling and random assignment serve different purposes.

3. Check all 44 observed counts

For this confidence interval, check observed successes and failures in each sample:

n1p^1=x1≥10andn1(1−p^1)=n1−x1≥10n_1\hat p_1=x_1\ge10\quad\text{and}\quad n_1(1-\hat p_1)=n_1-x_1\ge10
n2p^2=x2≥10andn2(1−p^2)=n2−x2≥10n_2\hat p_2=x_2\ge10\quad\text{and}\quad n_2(1-\hat p_2)=n_2-x_2\ge10
Visual guide 3: all 44 observed counts pass
Observed counts for the confidence-interval conditions.
GroupSuccesses xxFailures n−xn-xLarge-count result
School A124124200−124=76200-124=76Both ≥10\ge10
School B6969150−69=81150-69=81Both ≥10\ge10

For an interval, use observed successes and failures. Also justify the design and separate 10%10\% checks; a count table alone cannot establish those conditions.

All 44 counts pass. Saying “both sample sizes exceed 3030” is not enough: a sample of 100100 with only 33 successes fails this count criterion.

If a condition fails

  • Nonrandom selection: increasing nn does not repair the sampling design or guarantee population generalization.
  • Paired groups: do not apply the independent two-sample SE⁡\operatorname{SE} without addressing the dependence.
  • A 10%10\% check fails: the usual uncorrected SE⁡\operatorname{SE} is not automatically justified.
  • A success or failure count is below 1010: this standard zz-interval is not justified by the AP large-count criterion. Do not present its endpoints as a supported interval.
Quick check: do you use a hypothesized p0p_{0} for these count checks?

No. For this confidence interval, use each sample’s observed successes and failures. A hypothesized value belongs to a test setup, not these interval checks.

Build the confidence interval

For the main school example, construct a 95%95\% confidence interval for p1−p2p_{1}-p_{2}, the A-minus-B difference in population proportions preferring digital notes. The conditions have been justified above.

Step 1: calculate the point estimate

p^1−p^2=0.62−0.46=0.16\hat p_1-\hat p_2=0.62-0.46=0.16

Step 2: calculate the unpooled SE⁡\operatorname{SE}

SE⁡=0.62(0.38)200+0.46(0.54)150≈0.0532353\operatorname{SE}=\sqrt{\frac{0.62(0.38)}{200}+\frac{0.46(0.54)}{150}}\approx0.0532353

Step 3: choose the critical value z∗z^*

For a two-sided confidence level CC, place CC in the center of a standard normal curve. The remaining area, 1−C1-C, splits equally between the 22 tails. The positive boundary is z∗z^*.

Visual guide 4: split the unused area equally
Left tail0.0250.025

1−0.952\frac{1 – 0.95}{2}. The negative boundary is about −1.960-1.960.

Center0.950.95

The chosen 95%95\% confidence level determines the central standard normal area.

Right tail0.0250.025

The positive boundary is z∗≈1.960z^*\approx 1.960, with cumulative area 0.9750.975 to its left.

These cards show how a two-sided critical value is selected from the standard normal model. The 95%95\% is not a percentage of individual students within the interval.

Common two-sided standard normal critical values.
Confidence CCArea in each tailCumulative area for +z∗+z^*z∗z^* (rounded)
90%90\%0.0500.0500.9500.9501.6451.645
95%95\%0.0250.0250.9750.9751.9601.960
99%99\%0.0050.0050.9950.9952.5762.576

For 95%95\% confidence, C=0.95C=0.95 and each tail has area 0.0250.025. The cumulative area to the left of the positive boundary is 0.9750.975, so z∗≈1.960z^*\approx 1.960.

z∗=Φ−1 ⁣(1+C2)z^*=\Phi^{-1}\!\left(\frac{1+C}{2}\right)

Φ−1\Phi^{-1} gives the standard normal boundary at the stated cumulative area.

For example, use invNorm⁡(0.975,0,1)\operatorname{invNorm}(0.975,0,1) if your calculator follows the cumulative-area, mean, SD input order. Using invNorm⁡(0.95)\operatorname{invNorm}(0.95) would give the wrong critical value for a two-sided 95%95\% interval.

Step 4: multiply to get the margin of error

ME⁡=z∗SE⁡=z∗0.002834≈0.1043393\operatorname{ME}=z^*\operatorname{SE}=z^*\sqrt{0.002834}\approx0.1043393

Using rounded factors: ME⁡≈1.959964(0.0532353)≈0.1043393\operatorname{ME}\approx1.959964(0.0532353)\approx0.1043393.

The margin of error is about 10.4310.43 percentage points. It is the distance from the estimate to either endpoint. The full interval width is 22 times the margin of error.

Step 5: subtract and add the margin of error

Lower endpoint: 0.16−0.1043393≈0.05570.16-0.1043393\approx0.0557.

Upper endpoint: 0.16+0.1043393≈0.26430.16+0.1043393\approx0.2643.

95%95\% confidence interval for p1−p2p_1-p_2: (0.0557,0.2643)(0.0557,0.2643).

Visual guide 5: the estimate sits halfway between the endpoints
Construction of a confidence interval for a difference in population proportionsNumber line for a 95% confidence interval estimating the School A minus School B population preference-proportion difference. Sample results are 124/200 and 69/150. The point estimate is 16 percentage points, with margin of error about 10.43 percentage points in each direction. Endpoints are approximately 5.57 and 26.43 percentage points. Standard error is about 5.32 percentage points; z star is about 1.960. 0 8 16 24 32 A minus B (percentage points) −10.43 pp +10.43 pp 5.57 16.00 26.43 Point estimate: 16.00 pp Build the 95% interval around the sample difference A: 124/200 = .62 B: 69/150 = .46 Estimate: .16 95% CI: (5.57, 26.43) percentage points. SE ≈ 5.32 pp; z* ≈ 1.960. This is an interval on a number line, not a probability density curve. Construction of a confidence interval for a difference in population proportionsNumber line for a 95% confidence interval estimating the School A minus School B population preference-proportion difference. Sample results are 124/200 and 69/150. The point estimate is 16 percentage points, with margin of error about 10.43 percentage points in each direction. Endpoints are approximately 5.57 and 26.43 percentage points. Standard error is about 5.32 percentage points; z star is about 1.960. 0 8 16 24 32 A minus B (percentage points) −10.43 pp
Posted on Google Google
0000003998 : Abdad Alam Shamim Alam profile picture
0000003998 : Abdad Alam Shamim Alam
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
This is the bestest institution I have ever come to and I love it very much.
Posted on Google Google
Aliki S profile picture
Aliki S
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Ali profile picture
Ali
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Sly is him
Posted on Google Google
Jitendra Kumar Kumawat profile picture
Jitendra Kumar Kumawat
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Raahil Hasan profile picture
Raahil Hasan
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Tina Mudarres profile picture
Tina Mudarres
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
The classes are amazing and my child learnt so much
Posted on Google Google
alisha gadoya profile picture
alisha gadoya
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I have experienced a lot of good things, it has taught me so many things and from B/Cs I have gone to an A! This is wonderful and Mavish’s class is awesome.
Posted on Google Google
smasher 123 profile picture
smasher 123
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
It’s sensational my kid went. First he only would get C now it’s all A’s really good and recommend
Posted on Google Google
Mosa Al- Samaraie profile picture
Mosa Al- Samaraie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
AMAZING PRICES GREAT TEACHERS STRAIGHT FORWARD LEARNING STEADY PACE IN TUTORING
Posted on Google Google
Cael Dagnelie profile picture
Cael Dagnelie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I had a great experience learning math with Ms. Mavish. She explains complex topics in a very clear and simple way, which made it easier for me to understand and enjoy the subject. Her patience and dedication really stood out, and she always made sure that everyone in the class was keeping up. I especially appreciated how approachable she was — I never felt afraid to ask questions, and she was always willing to help. Thanks to her teaching, my confidence in math has grown a lot. I’m really thankful for the effort she puts into every lesson!

NUM8ERS is one of finest tutoring institutes in UAE, Located in Al Barsha 1, Dubai. Close to DUBAI AMERICAN ACADEMY (DAA) & AMERICAN SCHOOL OF DUBAI (ASD).