Constructing a Confidence Interval for the Difference Between Two Population Means
Estimate how far apart population averages are. Use independent samples to calculate a standard error, add a margin of error, and build a two-sample -interval with a clear meaning.
By the end of this lesson, you should be able to:
- Define the difference between population means in context.
- Choose a two-sample -interval and justify its conditions.
- Calculate the point estimate, standard error and margin of error.
- Use an appropriate critical value to construct the interval.
Before you start: Review the sampling distribution of in Topic 4.6, and the structure of a confidence interval.
First time learning this? Follow the delivery-time example, then practice the four steps: state, plan, calculate and conclude.
Here to revise? Use the formula checklist, then try the practice questions before revealing the solutions.
The concept in 60 seconds
sample averages give a point estimate of the difference between population averages. Because different samples would produce different estimates, report an interval to describe the uncertainty.
The difference in sample means is . The margin of error is a critical value times the estimated standard error of that difference.
The central idea: Center the interval on the observed sample difference. Extend it in both directions by an amount that reflects sample variability, sample sizes and the chosen confidence level.
This method estimates for independent groups when the population standard deviations are unknown. The interval concerns a difference between averages, not the difference between every pair of individuals.
How much longer does service A take?
A business wants to estimate the difference in population mean delivery times between services A and B. In this study, the population means and standard deviations are unknown. independent SRSs give these sample summaries:
| Group | Sample mean and | Sample and population sizes |
|---|---|---|
| 1: Service A | ; | from |
| 2: Service B | ; | from |
Let be the difference in mean delivery time, in minutes, for the defined delivery populations, using . A positive difference means A takes longer on average.
The point estimate is . A unpooled two-sample -interval will be approximately .
Pause and predict: Why not report exactly as the population difference? is the observed sample difference. The interval allows for sampling variation around that estimate.
We will justify the procedure, calculate its standard error and margin of error, and connect the endpoints to the original question.
Key ideas and notation
Parameter
is the difference between population means. Define both populations, the quantitative response and the subtraction order.
“A-minus-B mean delivery time” is more precise than “the difference.”
Point estimate
is the sample-based estimate at the center of the interval.
For the deliveries, it is .
Standard error
estimates how much the difference between sample means varies across repeated independent sample pairs.
Use each sample’s with its own sample size.
Margin of error
is the distance from the point estimate to either endpoint.
The full interval width is .
Confidence level and critical value
The chosen confidence level, such as , determines the central area used to find in a distribution with appropriate degrees of freedom. The confidence level describes the long-run success rate of the interval method under its conditions.
is a positive multiplier for interval construction. It is different from a test statistic calculated to test a null claim.
Choose the right method
Ask whether the response is quantitative and whether the groups are independent. Check that you are estimating a difference between population means and that population standard deviations are unknown.
Use a one-sample -interval for when appropriate. There is sample mean to estimate mean.
Form within-pair difference per complete pair. Use a one-sample -interval on the differences.
Use a two-sample -interval for , with separate sample means, standard deviations and sizes.
The study design chooses the method. Equal sample sizes do not create matched pairs, and different group labels do not automatically establish independence.
For the delivery study, the samples are independently selected from service populations, so the independent two-sample method fits the design.
Use the unpooled procedure here: Estimate each group’s contribution to variability separately. The formula in this lesson does not assume equal population variances. On the calculator, choose Pooled: No.
Check the conditions
A confidence interval needs more than summary statistics. Explain why the collection method and data support the procedure.
Apply the checks to the deliveries
- Randomization: The study gives independently selected SRSs.
- : and .
- Sample sizes: and support the course’s large-sample condition for the two-sample procedure.
These separate checks support treating observations within each sample as approximately independent. They do not make sampling without replacement exactly independent.
Small samples and experiments
When a sample is small, inspect the shape rather than adding the sample sizes together. Under the course’s small-sample check, both sample distributions should be free from strong skewness and outliers; approximately normal population models also support the method.
For a randomized experiment, explain the random assignment of treatments. The sampling-fraction condition is not required merely for treatment assignment; a separate finite-population sampling stage still needs its own check. The design and sample-data checks still matter.
If a check fails: State the limitation. A large sample does not remove volunteer bias, and a calculator cannot repair a paired design treated as independent.
Calculate the interval
Step 1: find the point estimate
Use the declared order: .
Step 2: calculate the standard error
Add the estimated variances of the separate sample means, then take the square root. The subtraction in the estimate does not mean subtracting standard deviations.
Step 3: find and the margin of error
For a unpooled interval, technology gives and . Thus:
Step 4: calculate both endpoints
Using the rounded margin of error, the interval is approximately , or .
The dot marks the observed A-minus-B sample mean difference. Each endpoint is margin of error from the estimate. The interval estimates a population mean difference, not a range of individual delivery times.
A concise interpretation: We are confident that service A’s population mean delivery time is between about and longer than service B’s for the defined delivery populations.
Rounding: Keep full precision for and while calculating. Round the final endpoints and include units. Using rounded displayed values can change the last decimal slightly.
critical values, degrees of freedom and precision
Population standard deviations are unknown, so the uses sample standard deviations. A critical value accounts for this additional uncertainty. For confidence level , expressed as a proportion with , leaves central area under the chosen curve and area in each tail.
For confidence, each tail has area . If you already know , the inverse- input is .
Why might be a decimal?
The usual unpooled two-sample procedure uses Welch’s approximate degrees of freedom, calculated from both sample sizes and standard deviations. Fractional is normal. Use the technology’s value without forcing it to .
The unpooled lies between the smaller of and and . For the deliveries, those bounds are and ; is within them.
Optional: where the technology’s comes from
Let and . Welch’s approximation is .
For the deliveries, and , giving . You can use a two-sample interval command to obtain this automatically.
If you are using a printed table
A conservative classroom option is . For the deliveries, gives and a interval of approximately .
This interval is slightly wider than the technology-based unpooled interval because the conservative gives a larger . State which method you used. If your table lacks the selected row, follow the table’s instructions; using an available lower preserves a conservative critical value.
What makes an interval wider or narrower?
- Higher confidence: a larger and wider interval for the same data.
- More variability: larger sample standard deviations increase and width, all else equal.
- Larger samples: reduce contributions and , improving precision when variability stays similar.
If both sample sizes are multiplied by and both sample standard deviations remain the same, becomes of its original value. Under the usual unpooled procedure, and