Sampling Distributions for the Difference Between Two Sample Means
Compare groups without losing track of sampling variation. Learn where differences between sample means are centered, how much they vary, and when a normal model can help calculate probabilities.
By the end of this lesson, you should be able to:
- Distinguish a difference between sample means from a difference between population means.
- Calculate the mean and standard deviation of .
- Justify randomization, independence and normal-model conditions for groups.
- Calculate and interpret a probability about repeated differences between sample means.
Before you start: Review sampling distributions for mean, variances of independent random variables, the central limit theorem and normal probabilities. Keep the group subtraction order fixed.
First time learning this? Follow the delivery-time example, then connect the center, spread and normal-model checks.
Here to revise? Use the formula checklist, then attempt the practice questions before opening the solutions.
The concept in 60 seconds
Take a random sample from each of groups and calculate their sample means. Subtract them in a stated order. That difference is an estimate of the difference between the population means.
If you repeat the sampling process, the sample means change. Their difference changes too. The distribution of all those possible differences is the sampling distribution of .
Three questions guide the model: Where is the distribution centered? How much do the differences vary? Is a normal shape justified?
Its center is . For independent sample means, add their variances to find the variance of the difference, then take a square root to obtain its standard deviation. Subtracting the means does not mean subtracting their variability.
Comparing delivery services
A business compares delivery times from services A and B. Treat each service’s defined set of deliveries as a separate population. For this teaching example, the population summaries are known:
| Group | Population mean and | Sample and population sizes |
|---|---|---|
| 1: Service A | ; | ; |
| 2: Service B | ; | ; |
Take the SRSs independently, without replacement within each population. Individual delivery times are right-skewed. Use the order , and define .
The population difference is . Across repeated pairs of samples, is centered at . It does not equal every time.
Pause and predict: Could a particular pair of samples give , even though the population difference is ? Yes. The sampling distribution helps us describe how often results that far above the center would occur.
For these sample sizes, has a standard deviation of approximately . Both samples are large enough for the course’s central-limit-theorem normal approximation, so we can estimate probabilities about .
Key ideas and notation
Population difference
is the difference between population means. It is a fixed parameter for the specified populations.
Delivery example: the A-minus-B population difference is .
Sample difference
is calculated from pair of samples. Before sampling, it is a random statistic.
is a shorter name for this difference; define its order before using it.
sources of variation
and describe individual values in the populations. and describe variation of the separate sample means.
Spread of the difference
describes variation of across repeated samples of the stated sizes.
Its units match the original measurement units.
Independent groups and matched pairs
separately selected groups can have independent sample means. Repeated measurements on the same people, or deliberately matched observations, are paired data. For pairs, analyze within-pair differences using their own mean and .
Having samples of equal size does not establish pairing. Having different group names does not establish independence; check how the data were collected.
Build the sampling distribution
Keep the populations and sample sizes fixed. Each repetition produces means and difference.
Use independent random selections from the same populations each time.
Calculate the average within each sample, then subtract A minus B.
A histogram of many repeated differences approximates the sampling distribution.
Reset to the same populations for each hypothetical repetition. A sampling-distribution axis contains differences of means, not individual delivery times.
Do not mix these distributions: The distribution of individual A delivery times, the distribution of , and the distribution of describe different quantities.
observed pair of sample means supplies point in this imagined distribution. You do not need to collect thousands of real samples to use the theoretical model; repeated sampling explains what the model represents.
Calculate the center and standard deviation
Center: subtract the population means
Each sample mean is centered at its own population mean, so their difference is centered at the population difference. For the deliveries, . This makes an unbiased estimator of under the stated sampling design.
Spread: add variances, then take the square root
Under the independent-observation model:
This variance rule requires independent sample means. For finite samples drawn without replacement, the calculations below use the unadjusted independent-observation approximation justified by the separate sampling checks. The subtraction sign changes the center; it does not cancel independent uncertainty. Squaring the coefficient gives in the variance calculation.
- For A’s sample mean: .
- For B’s sample mean: .
- Add the variances: .
- Take the square root: .
Interpretation: Across repeated independent random samples of A deliveries and B deliveries, the A-minus-B sample mean difference typically varies from the population difference of by about .
What changes when sample sizes grow?
Increasing either sample size decreases its contribution to the variance. Increasing both sample sizes by a factor of divides the total variance by and multiplies the standard deviation by . The center stays at .