Constructing a Confidence Interval for a Population Mean or Population Mean Difference
Turn a sample average into an interval estimate. Learn when to use , how to calculate a margin of error, and how paired observations become sample of differences.
By the end of this lesson, you should be able to:
- Choose a one-sample -interval and identify its population parameter.
- Describe -distributions and select using confidence level and degrees of freedom.
- Check the relevant conditions for single measurements or paired differences.
- Calculate standard error, margin of error and interval endpoints.
Before you start: Review sample means, sample , square roots, random sampling and Topic 4.1’s sampling distributions.
First time learning this? Follow the bottle example, then work through the paired-data example.
Here to revise? Review the key formulas, then try the practice questions.
The concept in 60 seconds
A sample mean is an estimate, not the whole answer. If a sample of bottles averages , we can use that value to estimate the population mean fill volume. A confidence interval adds a margin of error to show the uncertainty created by random sampling.
The construction: .
For a population mean with unknown population , use , after checking the conditions. The critical value comes from a -distribution with .
For paired data, first calculate difference for each pair. Then construct the same kind of one-sample interval using the mean and of those differences.
This lesson focuses on choosing and constructing the interval. We also practise a clear interpretation; evaluating claims from an interval is developed in Topic 4.3.
A bottle-filling investigation
Fictional teaching investigation: A quality-control team wants to estimate the mean fill volume of a lot of bottles. It selects an SRS of bottles without replacement. The sample mean is , and the sample standard deviation is . Previous process information supports an approximately normal population model. The population mean and population are unknown.
The team wants a confidence interval for , the mean fill volume of all bottles in this lot. It is estimating a mean quantitative measurement, not the proportion of bottles above a cutoff.
Predict: Would describe the uncertainty of an average properly? Which quantity measures individual variation, and which estimates variation in sample averages?
The sample average estimates the unknown population average.
The of individual values is ; the estimated of the sample mean is .
At confidence with , multiply by this critical value.
The confidence level and degrees of freedom determine . The sample statistics determine the center and estimated standard error. Check the sampling design and shape before interpreting the interval.
Topic 4.1 supplied population parameters to study sampling distributions. Here the population parameters are unknown, so we estimate sampling spread with and use rather than substituting automatically.
Key ideas and notation
| Symbol / term | Meaning | In the bottle investigation |
|---|---|---|
| The population mean, a fixed unknown parameter. | Mean fill volume of all bottles in this lot. | |
| The sample mean, our point estimate of . | . | |
| and | Population and sample ; is unknown here. | Use to estimate individual population variability. |
| and | Number sampled and number in the population. | bottles; bottles. |
| Estimated standard deviation of the sample mean: . | . | |
| Positive critical value for the central confidence area of the relevant -distribution. | About for and . | |
| Degrees of freedom for a one-sample procedure. | . | |
| Margin of error: . | About . | |
| Population mean of the paired differences, in a stated order. | For completion times, the mean time reduction in the population. | |
| and | Mean and sample of the observed paired differences. | Compute them from the difference list, not from either original list alone. |
Why ? Once the sample mean is calculated, the deviations from that mean must sum to . Only of those deviations can vary freely. That is the degrees of freedom used in the one-sample model.
Understand and
A -distribution is symmetric and bell-shaped around . Compared with the standard normal curve, it has heavier tails, especially at small degrees of freedom. It allows for the extra uncertainty caused by estimating with .
The family contains different curves for different values. As increases, the curves approach the standard normal curve. Do not claim that every curve has . The standardized model is not an ordinary normal distribution.
For a two-sided interval, the central area is and each tail has area . The positive cutoff therefore has cumulative area to its left.