AP Statistics / Unit 3: Inference for Categorical Data: Proportions / Topic 3.9
NUM8ERS study notes · Topic 3.9

Sampling Distributions for the Difference Between Sample Proportions

22 samples give 22 proportions. Subtract them to compare the groups, then use a sampling distribution to understand how that difference changes from sample to sample.

2026–27 curriculum6 worked examples10 practice questions6 visual guides

By the end of this lesson, you should be able to:

  • Define the 22 population proportions and keep a consistent subtraction order.
  • Distinguish an observed sample difference from its sampling distribution.
  • Calculate and interpret the distribution’s mean and standard deviation.
  • Check the random-sampling design, independence, both 10%10\% conditions and all 44 expected counts.
  • Use a justified normal approximation to find a probability for a sample difference.
  • Explain what larger samples and reversed group order change.

Before you start: Review Topic 3.2’s sampling distribution for 11 proportion. Be comfortable with p^=xn\hat{p} =\frac{x}{n}, standard deviation, zz-scores and normal-curve areas.

First time learning this? Follow the school example through the notation, conditions and probability calculation.

Here to revise? Use the sampling-model checklist, then try the questions before opening hints or solutions.

The concept in 60 seconds

Suppose we compare the proportion of students who prefer digital study notes at School A with the corresponding proportion at School B. 11 random sample comes from each school.

D=p^1−p^2D=\hat p_1-\hat p_2

Read this as “the sample proportion from group 1 minus the sample proportion from group 2.” DD is a short name for that difference.

Different random samples will usually give different values of DD. The sampling distribution of p^1−p^2\hat{p}_{1} – \hat{p}_{2} describes those possible differences and how often they would occur across repeated independent sampling with the same sample sizes.

Under the appropriate independent-sampling model:

  • Center: the mean sample difference equals the true population difference, p1−p2p_{1} – p_{2}.
  • Spread: both samples contribute variability. Add their variances, then take the square root.
  • Shape: an approximately normal model is justified when all 44 expected success/failure counts are at least 1010, with the sampling conditions also satisfied.

11 sample pair gives 11 statistic. A sampling distribution describes the statistic across many possible sample pairs. It is not a distribution of individual students’ yes/no responses.

Quick check: does 11 observed difference tell us the exact population difference?

No. A sample difference varies around the population difference. It is an estimator, not an exact population measurement. A confidence interval will use this variability in Topic 3.10.

22 groups, 11 difference

All school examples here are fictional. For the main probability model, suppose the true preference proportions are p1=0.60p_{1} = 0.60 at School A and p2=0.45p_{2} = 0.45 at School B. These population values are stipulated for learning; they are not established by an observed sample.

Take an independent SRS of 200200 students from School A’s 10,00010{,}000 students and an independent SRS of 150150 from School B’s 8,0008{,}000 students, without replacement. Use the same yes/no preference question in both schools.

Visual guide 1: 22 independent samples produce 11 difference
Group 1 · School Ap^1=x1200\hat{p}_{1} =\frac{x_{1}}{200}

Count digital-notes preferences in an SRS of 200200 School A students.

Group 2 · School Bp^2=x2150\hat{p}_{2} =\frac{x_{2}}{150}

Independently count digital-notes preferences in an SRS of 150150 School B students.

Combine in this orderD=p^1−p^2D = \hat{p}_{1} – \hat{p}_{2}

Each sample pair contributes 11 A-minus-B difference. Repeat conceptually to form its sampling distribution.

Use the same definition of “prefers digital notes” in both populations. The group order remains fixed for every sample pair. These cards show a study process, not results from an actual survey.

Population values stay fixed; sample values vary.
SymbolWhat it meansSchool model
p1p_{1}, p2p_{2}True population success proportions0.600.60 at School A; 0.450.45 at School B
n1n_{1}, n2n_{2}Numbers sampled from each group200200 from A; 150150 from B
N1N_{1}, N2N_{2}Sizes of the respective populations10,00010{,}000 at A; 8,0008{,}000 at B
x1x_{1}, x2x_{2}Success counts in 11 sample pairHypothetical observed counts 124124 and 6969
p^1\hat{p}_{1}, p^2\hat{p}_{2}Sample proportions x1n1\frac{x_{1}}{n_{1}} and x2n2\frac{x_{2}}{n_{2}}Hypothetical observed values 0.620.62 and 0.460.46
p1−p2p_{1} – p_{2}Population difference: the parameter0.150.15 in the stipulated model
D=p^1−p^2D = \hat{p}_{1} – \hat{p}_{2}Sample difference: the statistic0.160.16 for the hypothetical sample pair

Compute 11 observed difference

For a hypothetical sample pair, 124124 of the 200200 School A students and 6969 of the 150150 School B students prefer digital notes.

p^1=124200=0.62\hat p_1=\frac{124}{200}=0.62
p^2=69150=0.46\hat p_2=\frac{69}{150}=0.46
D=0.62−0.46=0.16D=0.62-0.46=0.16

The sample preference proportion is 1616 percentage points higher at School A. The true difference stipulated for the model is 0.60−0.45=0.150.60 – 0.45 = 0.15, or 1515 percentage points. A sample estimate can differ from the true value.

Do not call 0.160.16 a “16%16\% relative increase.” A subtraction of proportions gives a percentage-point difference when multiplied by 100100. A relative percentage change would require a separate denominator.

Visual guide 2: fixed parameter versus variable statistic
Population differencep1−p2=0.15p_{1} – p_{2} = 0.15

A fixed A-minus-B preference difference of 1515 percentage points, stipulated for the model.

11 sample differencep^1−p^2=0.16\hat{p}_{1} – \hat{p}_{2} = 0.16

An estimate from the illustrative 124200\frac{124}{200} and 69150\frac{69}{150} sample results. Another sample pair could give another value.

The sampling distribution describes repeated values of the statistic around the fixed parameter. These values are teaching examples, not observations from a real school study.

What is repeated?

  1. Draw a new SRS of 200200 from School A and independently draw a new SRS of 150150 from School B.
  2. Calculate the 22 new sample proportions.
  3. Subtract in the same A-minus-B order.
  4. Repeat the process conceptually to build the distribution of DD.

The populations and their true proportions stay fixed. The selected students and sample results change. Repeated sampling is a model for understanding uncertainty; you do not need to collect thousands of actual samples.

Quick check: why can the 22 samples have different sizes?

Equal sizes are not required. Calculate each proportion using its own denominator, and keep its nn matched to its pp in the variance formula. Here, n1=200n_{1} = 200 and n2=150n_{2} = 150.

Mean and standard deviation

The mean

μD=p1−p2\mu_D=p_1-p_2

Main model:

μD=0.60−0.45=0.15\mu_D=0.60-0.45=0.15

Across repeated independent random samples of these sizes, the average sample preference difference, School A minus School B, equals the population difference of 0.150.15. In this model, p^1−p^2\hat{p}_{1} – \hat{p}_{2} is an unbiased estimator of p1−p2p_{1} – p_{2}.

The mean is not automatically 00. It is 00 when the population proportions are equal. Here, the specified populations differ, so the model is centered at 0.150.15.

The standard deviation

σD=p1(1−p1)n1+p2(1−p2)n2\sigma_D=\sqrt{\frac{p_1(1-p_1)}{n_1}+\frac{p_2(1-p_2)}{n_2}}

Use the population proportions stated in the sampling model and each group’s own sample size.

Visual guide 3: add the uncertainty from both groups
School A variance0.001200.00120

p1(1−p1)n1=0.60(0.40)200\frac{p_{1}(1 – p_{1})}{n_{1}}=\frac{0.60(0.40)}{200}.

School B variance0.001650.00165

p2(1−p2)n2=0.45(0.55)150\frac{p_{2}(1 – p_{2})}{n_{2}}=\frac{0.45(0.55)}{150}.

Difference SD0.00285≈0.05339\sqrt{0.00285} \approx 0.05339

Add 0.00120+0.001650.00120 + 0.00165, then take the square root. Independence is required for this addition rule.

Means subtract: 0.60−0.45=0.150.60 – 0.45 = 0.15. Independent variance contributions add. Variance is in squared-proportion units; SD returns to proportion-difference units.

Under the independent-observation approximation, the main school model gives:

Var⁡(p^1)=0.60(0.40)200=0.00120\operatorname{Var}(\hat p_1)=\frac{0.60(0.40)}{200}=0.00120
Var⁡(p^2)=0.45(0.55)150=0.00165\operatorname{Var}(\hat p_2)=\frac{0.45(0.55)}{150}=0.00165
Var⁡(D)=0.00120+0.00165=0.00285\operatorname{Var}(D)=0.00120+0.00165=0.00285
σD=0.00285≈0.0533854\sigma_D=\sqrt{0.00285}\approx0.0533854

The sample differences typically vary by about 0.05340.0534, or 5.345.34 percentage points, from the true difference of 0.150.15. This is a measure of sampling spread, not a guaranteed error bound.

Why is there a plus sign inside the square root?

Either sample can make the difference fluctuate. Subtracting the second proportion changes its direction but does not cancel its variability. For independent random variables, the variance of a difference is the sum of their variances.

Do not subtract the 22 standard deviations, and do not simply add them. The correct route is square → add the variances → square root.

Optional explanation: where does independence enter the variance formula?

In general, Var⁡(X−Y)=Var⁡(X)+Var⁡(Y)−2Cov⁡(X,Y)\operatorname{Var}(X-Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)-2\operatorname{Cov}(X,Y). Independent variables have covariance 00, giving the sum of variances. The simpler AP two-proportion formula therefore depends on an appropriate independent-group design.

For example, surveying the same students before and after a change creates paired responses. That dependence is not handled by plugging 22 marginal proportions into the independent-samples formula.

For SRSs without replacement from finite populations, the displayed SD is an approximation that omits finite-population correction factors. The separate 10%10\% checks below justify this usual AP approximation. Under independent Bernoulli sampling, the variance expression itself is exact.

Quick check: if the 22 population proportions are equal, must every sample difference be 00?

No. The distribution’s mean is 00, but individual sample differences still vary. A center of 00 does not mean a standard deviation of 00.

Check the conditions

Before using the SD formula or a normal probability, explain how the data were collected and check the relevant conditions. Writing “the sample is large” is not enough.

1. Randomization and independent groups

For a sampling study, use 22 independent random samples. The sampling process for 11 school must not determine the selected students or responses in the other. The main example explicitly specifies independent SRSs.

A large voluntary online poll does not meet the random-sampling condition merely because it has many responses. Similarly, the same people measured 22 times are paired, not 22 independent samples.

For an appropriate two-group experiment, randomly assign treatments to experimental units. Random assignment supports a treatment comparison; random sampling supports generalization to the sampled populations. The finite-population 10%10\% sampling condition is not required merely because subjects were assigned to treatment groups. Do not confuse “randomly assigned” with “randomly selected.”

2. A separate 10%10\% condition for each sampled population

When sampling without replacement, verify:

n1≤0.10N1n_1\le0.10N_1 and n2≤0.10N2n_2\le0.10N_2.

School A: 200≤0.10(10,000)=1,000200\le0.10(10{,}000)=1{,}000.

School B: 150≤0.10(8,000)=800150\le0.10(8{,}000)=800.

Both pass. These checks let us treat observations within each sample as approximately independent for the usual variance approximation. Checking n1+n2(N1+N2)\frac{n_{1} + n_{2}}{(N_{1} + N_{2})} is not a substitute; each sample must fit its own population.

3. All 44 large-count checks for normality

For the theoretical sampling distribution, use the specified true population proportions:

n1p1≥10andn1(1−p1)≥10n_1p_1\ge10\quad\text{and}\quad n_1(1-p_1)\ge10
n2p2≥10andn2(1−p2)≥10n_2p_2\ge10\quad\text{and}\quad n_2(1-p_2)\ge10
Visual guide 4: all 44 expected counts must pass
Normality checks for the stipulated true proportions.
GroupExpected successesExpected failuresResult
School A200(0.60)=120200(0.60) = 120200(0.40)=80200(0.40) = 80Both ≥10\ge10
School B150(0.45)=67.5150(0.45) = 67.5150(0.55)=82.5150(0.55) = 82.5Both ≥10\ge10

These counts justify approximate normal shape after the sampling conditions are established. They use population pp values, not pooled proportions. The main example separately satisfies 200≤1,000200 \le 1{,}000 and 150≤800150 \le 800 for sampling without replacement.

Expected counts can be nonintegers. The model expects 67.567.5 successes on average in the School B sample; an actual sample cannot have 12\frac12 of a student. A normal shape for the difference is justified by the 44 expected counts, not by requiring n1+n2≥30n_{1} + n_{2} \ge 30.

If a condition fails

  • No appropriate random design: a normal calculation does not repair selection bias or establish the intended generalization.
  • Paired or dependent groups: the independent-group variance formula needs a different model.
  • A 10%10\% check fails: do not automatically use the uncorrected independent-observation SD.
  • An expected count is below 1010: the usual normal approximation is not justified by this AP criterion. The statistic still has a sampling distribution, but this normal method is not supported.

Separate the jobs. Randomization addresses design. The 10%10\% checks address within-sample dependence when sampling without replacement. The 44 count checks address approximate normal shape. One check cannot replace another.

Quick check: if group 1 has 200200 expected successes but group 2 has only 44, does the normality check pass?

No. Each group needs at least 1010 expected successes and at least 1010 expected failures. Large counts in 11 group cannot compensate for small counts in the other.

Normal probabilities for the sample difference

For the justified school model, D=p^1−p^2D = \hat{p}_{1} – \hat{p}_{2} is approximately normal with mean 0.150.15 and SD 0.05338540.0533854. State both quantities rather than writing N(0.15,0.00285)N(0.15, 0.00285) without saying whether the second value is a variance or an SD.

z=d−(p1−p2)σDz=\frac{d-(p_1-p_2)}{\sigma_D}

dd is the sample-difference boundary in the probability question.

Standardize the difference as 11 statistic. Do not calculate 22 unrelated tail probabilities and subtract them.

Visual guide 5: turn sample-difference events into areas
Two probability events under the same sample-proportion difference distributionTwo panels show the same normal approximation for D, the School A minus School B sample-proportion difference. Both use mean 0.15 and SD about 0.0533854, for stipulated p1 0.60, p2 0.45, n1 200 and n2 150. The upper panel shades the right tail at or above 0.25, area about 0.0305. The lower panel shades 0.10 to 0.20, area about 0.6510. Dotted lines locate the true population difference and dashed lines locate probability bounds. 0 5 10 Right area ≈ 0.0305 Difference at least .25 −.05 0 .10 .15 .20 .25 .35 Sample difference D = p̂₁ − p̂₂ 0 5 10 Central area ≈ 0.6510 Difference between .10 and .20
Posted on Google Google
0000003998 : Abdad Alam Shamim Alam profile picture
0000003998 : Abdad Alam Shamim Alam
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
This is the bestest institution I have ever come to and I love it very much.
Posted on Google Google
Aliki S profile picture
Aliki S
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Ali profile picture
Ali
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Sly is him
Posted on Google Google
Jitendra Kumar Kumawat profile picture
Jitendra Kumar Kumawat
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Raahil Hasan profile picture
Raahil Hasan
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Tina Mudarres profile picture
Tina Mudarres
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
The classes are amazing and my child learnt so much
Posted on Google Google
alisha gadoya profile picture
alisha gadoya
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I have experienced a lot of good things, it has taught me so many things and from B/Cs I have gone to an A! This is wonderful and Mavish’s class is awesome.
Posted on Google Google
smasher 123 profile picture
smasher 123
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
It’s sensational my kid went. First he only would get C now it’s all A’s really good and recommend
Posted on Google Google
Mosa Al- Samaraie profile picture
Mosa Al- Samaraie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
AMAZING PRICES GREAT TEACHERS STRAIGHT FORWARD LEARNING STEADY PACE IN TUTORING
Posted on Google Google
Cael Dagnelie profile picture
Cael Dagnelie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I had a great experience learning math with Ms. Mavish. She explains complex topics in a very clear and simple way, which made it easier for me to understand and enjoy the subject. Her patience and dedication really stood out, and she always made sure that everyone in the class was keeping up. I especially appreciated how approachable she was — I never felt afraid to ask questions, and she was always willing to help. Thanks to her teaching, my confidence in math has grown a lot. I’m really thankful for the effort she puts into every lesson!

NUM8ERS is one of finest tutoring institutes in UAE, Located in Al Barsha 1, Dubai. Close to DUBAI AMERICAN ACADEMY (DAA) & AMERICAN SCHOOL OF DUBAI (ASD).