Sampling Distributions for the Difference Between Sample Proportions
samples give proportions. Subtract them to compare the groups, then use a sampling distribution to understand how that difference changes from sample to sample.
By the end of this lesson, you should be able to:
- Define the population proportions and keep a consistent subtraction order.
- Distinguish an observed sample difference from its sampling distribution.
- Calculate and interpret the distribution’s mean and standard deviation.
- Check the random-sampling design, independence, both conditions and all expected counts.
- Use a justified normal approximation to find a probability for a sample difference.
- Explain what larger samples and reversed group order change.
Before you start: Review Topic 3.2’s sampling distribution for proportion. Be comfortable with , standard deviation, -scores and normal-curve areas.
First time learning this? Follow the school example through the notation, conditions and probability calculation.
Here to revise? Use the sampling-model checklist, then try the questions before opening hints or solutions.
The concept in 60 seconds
Suppose we compare the proportion of students who prefer digital study notes at School A with the corresponding proportion at School B. random sample comes from each school.
Read this as “the sample proportion from group 1 minus the sample proportion from group 2.” is a short name for that difference.
Different random samples will usually give different values of . The sampling distribution of describes those possible differences and how often they would occur across repeated independent sampling with the same sample sizes.
Under the appropriate independent-sampling model:
- Center: the mean sample difference equals the true population difference, .
- Spread: both samples contribute variability. Add their variances, then take the square root.
- Shape: an approximately normal model is justified when all expected success/failure counts are at least , with the sampling conditions also satisfied.
sample pair gives statistic. A sampling distribution describes the statistic across many possible sample pairs. It is not a distribution of individual students’ yes/no responses.
Quick check: does observed difference tell us the exact population difference?
No. A sample difference varies around the population difference. It is an estimator, not an exact population measurement. A confidence interval will use this variability in Topic 3.10.
groups, difference
All school examples here are fictional. For the main probability model, suppose the true preference proportions are at School A and at School B. These population values are stipulated for learning; they are not established by an observed sample.
Take an independent SRS of students from School A’s students and an independent SRS of from School B’s students, without replacement. Use the same yes/no preference question in both schools.
Count digital-notes preferences in an SRS of School A students.
Independently count digital-notes preferences in an SRS of School B students.
Each sample pair contributes A-minus-B difference. Repeat conceptually to form its sampling distribution.
Use the same definition of “prefers digital notes” in both populations. The group order remains fixed for every sample pair. These cards show a study process, not results from an actual survey.
| Symbol | What it means | School model |
|---|---|---|
| , | True population success proportions | at School A; at School B |
| , | Numbers sampled from each group | from A; from B |
| , | Sizes of the respective populations | at A; at B |
| , | Success counts in sample pair | Hypothetical observed counts and |
| , | Sample proportions and | Hypothetical observed values and |
| Population difference: the parameter | in the stipulated model | |
| Sample difference: the statistic | for the hypothetical sample pair |
Compute observed difference
For a hypothetical sample pair, of the School A students and of the School B students prefer digital notes.
The sample preference proportion is percentage points higher at School A. The true difference stipulated for the model is , or percentage points. A sample estimate can differ from the true value.
Do not call a “ relative increase.” A subtraction of proportions gives a percentage-point difference when multiplied by . A relative percentage change would require a separate denominator.
A fixed A-minus-B preference difference of percentage points, stipulated for the model.
An estimate from the illustrative and sample results. Another sample pair could give another value.
The sampling distribution describes repeated values of the statistic around the fixed parameter. These values are teaching examples, not observations from a real school study.
What is repeated?
- Draw a new SRS of from School A and independently draw a new SRS of from School B.
- Calculate the new sample proportions.
- Subtract in the same A-minus-B order.
- Repeat the process conceptually to build the distribution of .
The populations and their true proportions stay fixed. The selected students and sample results change. Repeated sampling is a model for understanding uncertainty; you do not need to collect thousands of actual samples.
Quick check: why can the samples have different sizes?
Equal sizes are not required. Calculate each proportion using its own denominator, and keep its matched to its in the variance formula. Here, and .
Mean and standard deviation
The mean
Main model:
Across repeated independent random samples of these sizes, the average sample preference difference, School A minus School B, equals the population difference of . In this model, is an unbiased estimator of .
The mean is not automatically . It is when the population proportions are equal. Here, the specified populations differ, so the model is centered at .
The standard deviation
Use the population proportions stated in the sampling model and each group’s own sample size.
.
.
Add , then take the square root. Independence is required for this addition rule.
Means subtract: . Independent variance contributions add. Variance is in squared-proportion units; SD returns to proportion-difference units.
Under the independent-observation approximation, the main school model gives:
The sample differences typically vary by about , or percentage points, from the true difference of . This is a measure of sampling spread, not a guaranteed error bound.
Why is there a plus sign inside the square root?
Either sample can make the difference fluctuate. Subtracting the second proportion changes its direction but does not cancel its variability. For independent random variables, the variance of a difference is the sum of their variances.
Do not subtract the standard deviations, and do not simply add them. The correct route is square → add the variances → square root.
Optional explanation: where does independence enter the variance formula?
In general, . Independent variables have covariance , giving the sum of variances. The simpler AP two-proportion formula therefore depends on an appropriate independent-group design.
For example, surveying the same students before and after a change creates paired responses. That dependence is not handled by plugging marginal proportions into the independent-samples formula.
For SRSs without replacement from finite populations, the displayed SD is an approximation that omits finite-population correction factors. The separate checks below justify this usual AP approximation. Under independent Bernoulli sampling, the variance expression itself is exact.
Quick check: if the population proportions are equal, must every sample difference be ?
No. The distribution’s mean is , but individual sample differences still vary. A center of does not mean a standard deviation of .
Check the conditions
Before using the SD formula or a normal probability, explain how the data were collected and check the relevant conditions. Writing “the sample is large” is not enough.
1. Randomization and independent groups
For a sampling study, use independent random samples. The sampling process for school must not determine the selected students or responses in the other. The main example explicitly specifies independent SRSs.
A large voluntary online poll does not meet the random-sampling condition merely because it has many responses. Similarly, the same people measured times are paired, not independent samples.
For an appropriate two-group experiment, randomly assign treatments to experimental units. Random assignment supports a treatment comparison; random sampling supports generalization to the sampled populations. The finite-population sampling condition is not required merely because subjects were assigned to treatment groups. Do not confuse “randomly assigned” with “randomly selected.”
2. A separate condition for each sampled population
When sampling without replacement, verify:
and .
School A: .
School B: .
Both pass. These checks let us treat observations within each sample as approximately independent for the usual variance approximation. Checking is not a substitute; each sample must fit its own population.
3. All large-count checks for normality
For the theoretical sampling distribution, use the specified true population proportions:
| Group | Expected successes | Expected failures | Result |
|---|---|---|---|
| School A | Both | ||
| School B | Both |
These counts justify approximate normal shape after the sampling conditions are established. They use population values, not pooled proportions. The main example separately satisfies and for sampling without replacement.
Expected counts can be nonintegers. The model expects successes on average in the School B sample; an actual sample cannot have of a student. A normal shape for the difference is justified by the expected counts, not by requiring .
If a condition fails
- No appropriate random design: a normal calculation does not repair selection bias or establish the intended generalization.
- Paired or dependent groups: the independent-group variance formula needs a different model.
- A check fails: do not automatically use the uncorrected independent-observation SD.
- An expected count is below : the usual normal approximation is not justified by this AP criterion. The statistic still has a sampling distribution, but this normal method is not supported.
Separate the jobs. Randomization addresses design. The checks address within-sample dependence when sampling without replacement. The count checks address approximate normal shape. One check cannot replace another.
Quick check: if group 1 has expected successes but group 2 has only , does the normality check pass?
No. Each group needs at least expected successes and at least expected failures. Large counts in group cannot compensate for small counts in the other.
Normal probabilities for the sample difference
For the justified school model, is approximately normal with mean and SD . State both quantities rather than writing without saying whether the second value is a variance or an SD.
is the sample-difference boundary in the probability question.
Standardize the difference as statistic. Do not calculate unrelated tail probabilities and subtract them.