Sampling Distributions for Sample Means
Why are averages more consistent than individual measurements? Follow a bottle-filling example to understand the center, spread and shape of sample means—and use that model to calculate probabilities.
By the end of this lesson, you should be able to:
- Calculate and interpret the center and spread of a sample-mean distribution.
- Justify randomization, the condition and a normal approximation.
- Find a probability for a sample average and explain it in context.
- Describe how changing sample size affects variability.
Before you start: Review means and standard deviations, square roots, -scores and the central limit theorem.
First time learning this? Start with the bottle example, then check the conditions before the worked examples.
Here to revise? Review the key formulas, then try the practice questions.
The concept in 60 seconds
A sample mean changes from sample to sample. Choose bottles at random and average their fill volumes. Choose another and calculate another average. The averages will usually differ, even though they come from the same population.
The sampling distribution of the sample mean describes the possible averages from all random samples of fixed size, together with how likely those averages are.
Remember three questions: Where are the averages centered? How much do they vary? What shape does their distribution have?
Under appropriate independence conditions, the center is and the standard deviation is . A normal population, or a sufficiently large sample, lets us use a normal model for the averages.
The key change from proportions is the variable: we now average a quantitative measurement, such as volume, mass or time. We are not counting a success category.
A bottle-filling example
Fictional teaching scenario: A production lot contains bottles. Treat the distribution of their fill volumes as approximately normal, with population mean and population standard deviation . Quality control takes a simple random sample of bottles without replacement and calculates its mean fill volume.
The population parameters are supplied for this model exercise. In real investigations, they are often unknown and must be estimated.
Predict: Would bottle’s volume or the average of bottles usually be closer to ? Explain before looking at the graph.
Select bottles from the same lot. Record individual volumes.
Add the volumes and divide by . This produces value of .
Repeat with fresh random samples of the same size. The collection of means illustrates the sampling distribution.
A sampling distribution is theoretical: it includes all possible samples under the stated design. A simulation using many repeated samples approximates it; a graph of the original volumes is a different distribution.
For observed sample, might be . That is sample result, not the mean of the entire sampling distribution. Across repeated samples, the sample means are centered at .
Key ideas and notation
| Symbol / idea | Meaning | Bottle example |
|---|---|---|
| A random individual quantitative measurement. | randomly selected bottle’s volume, in mL. | |
| The fixed population mean. | across the lot. | |
| The population standard deviation of individual measurements. | for individual bottle volumes. | |
| (x-bar) | The mean of the values in a random sample; before sampling it varies. | The average volume of randomly selected bottles. |
| and | Sample size and population size. | bottles; bottles. |
| The mean of the sampling distribution of . | : the long-run center of sample means. | |
| The standard deviation of the sampling distribution of . | , using the independence approximation. | |
| and | Sample and estimated standard deviation of . | If is unknown, estimate the spread with . |
One data set versus many possible data sets: The sample data distribution has individual values. The sampling distribution has possible values of a statistic, mean per possible sample. Both use mL here, but their horizontal axes represent different things.
Calling an unbiased estimator means its sampling distribution is centered at under the random sampling model. It does not mean that every sample mean equals .
Center, spread and sample size
Center:
Spread: , for independent observations.
The first formula says averaging preserves the population center. The second says averaging reduces the variability of the resulting statistic. Neither formula requires a normal population. Normality matters when we use a normal curve to find probabilities.
When an SRS is taken without replacement, the sampled values are not exactly independent. If , AP Statistics uses as an appropriate approximation. For the bottle lot, , so the approximation is justified. With a large sampling fraction, the exact without-replacement spread needs a finite-population adjustment; do not apply automatically.
| Sample size | Center | Spread |
|---|---|---|