Sampling Distributions for Sample Proportions
Different random samples give different proportions. Learn where those proportions are centered, how much they vary, when a normal curve is appropriate, and how to calculate and explain a probability about a sample.
By the end of this lesson, you should be able to:
- Distinguish a population proportion from a sample proportion .
- Calculate the mean and standard deviation of under a suitable sampling model.
- Check random sampling, the condition and expected success/failure counts.
- Describe how sample size affects the spread and shape of ‘s sampling distribution.
- Calculate normal-model probabilities and interpret them in the sampling context.
Before you start: Know proportions, sampling distributions and normal probabilities. Review Topic 3.1: Estimators for unbiasedness and Topic 2.11: The Normal Distribution for z-scores and areas.
First time learning this? Follow the school survey, check its conditions, then work through the shaded probability example.
Here to revise? Use the formula checklist, then try the practice questions before revealing solutions.
The concept in 60 seconds
Imagine a school of students where prefer an earlier lunch. The population proportion is . Take a simple random sample of students without replacement. If say yes, the sample proportion is .
A fresh random sample of might give or another value. The sampling distribution of describes all possible sample proportions and how likely they are under the specified sampling method and sample size.
Each student gives yes/no response.
Suppose say yes.
Calculate .
This sample contributes plotted value.
Take fresh samples of the same size and record each time.
The graph describes how those answers vary.
Keep the population, sampling rule and fixed while constructing one sampling distribution. A histogram of sample’s raw binary responses is a different graph.
Remember the target: stays fixed for this population. varies from sample to sample. A graph of repeated values is a graph of sample summaries, not a graph of individual students’ yes/no responses.
The proportions tend to cluster around . Larger random samples usually give a tighter cluster. With enough expected successes and failures, we can use an approximately normal sampling model.
All scenarios here are fictional teaching examples. Population proportions are supplied so we can study sampling behavior. The shape comparison uses an actual seeded simulation from a specified model; it is not collected school data.
Quick check: a sample of students gives how many values?
. Calculate the number saying yes divided by . To plot sample proportions, repeat the sampling process times and record proportion each time.
Notation and the sampling model
| Symbol | Meaning | School example |
|---|---|---|
| Population proportion in the counted category | ||
| Proportion in sample; | ||
| Success count in sample | students | |
| Number of observations per sample | students | |
| Size of the finite population | students |
A “success” is the category being counted. It might mean preferring an earlier lunch, owning a bicycle, or receiving a defective item. The word does not imply that the outcome is desirable.
School sample: .
School population: .
Connect a count to a proportion
For independent binary observations with the same success probability , the success count is , and . The binomial count and the sample proportion describe the same sample on different scales.
For , possible proportions are . The true sampling distribution is discrete. A normal curve is a continuous approximation to those possible values.
In a finite population sampled without replacement, the observations are dependent. A small sampling fraction lets us approximate the process with independent draws. If the fraction is substantial, use a model that accounts for sampling without replacement; the binomial model is then not exact.
Quick check: with , what count corresponds to ?
. The event instead corresponds to . Including a boundary matters for the exact discrete distribution, even though a continuous normal curve assigns a single point area.
Center, spread and sample size
For independent observations with the same population success probability :
Under the independent model:
The mean equals , so is an unbiased estimator under this model. The SD measures how much sample proportions vary across repeated samples. These formulas do not require a normal shape; normality is a separate question.
Apply the formulas to the school
Here and . The sampling fraction is small enough to support the independent approximation, as we will check below.
Mean is . The school’s sampling SD is approximately
The model is centered at , with a sampling SD of about percentage points. It is not centered at the observed . We use the supplied to describe the model for repeated samples.
What changes when increases?
Both are theoretical normal approximations for the same ; they are not simulation histograms. The shared horizontal range is . Density height is not a probability; probability is area under a curve. The full normal curves each have total area .
For , the approximate SD is , or percentage points. The center remains . Multiplying by halves the SD because is inside a square root.
Increasing sample size changes the statistic’s sampling distribution. Increasing the number of simulation repetitions at fixed only gives a more stable picture of that same distribution.
Do not swap for automatically: This lesson describes a sampling distribution using a supplied population proportion . Estimating its SD from sample data uses a standard error involving ; that is developed with confidence intervals in Topic 3.3.
Quick check: if doubles, does the SD halve?
No. It is multiplied by . To halve the SD while keeping fixed under the same independent model, multiply by .
Check the conditions
Ask: How were the observations selected?
A simple random sample supports inference to the intended population.
When sampling without replacement: .
A small sampling fraction supports the usual independent-model SD approximation.
Check both: and .
This supports an approximately normal sampling distribution.
Passing one check does not establish the others. Use from the supplied population model for the expected counts in this lesson.
1. Random sampling
The data should come from an appropriate random sample of the target population. In the school example, the sample is a simple random sample. A large volunteer sample would not satisfy this selection requirement simply because it contains many people.
2. The condition when sampling without replacement
School check: . The condition is met.
This condition supports treating the dependence from sampling without replacement as small when using . It does not make the observations exactly independent, and it does not establish a normal shape.
For genuinely independent sampling, such as a specified independent binary-response model or sampling with replacement, this finite-population check is not needed. Independence must still be justified by the process.
3. Enough expected successes and failures
School check: expected successes; expected failures. Both are at least .
These counts support an approximately normal shape. Use the supplied population proportion . The observed sample may contain successes, but the expected count under is .
A full justification for the school: “The students form a simple random sample. Sampling without replacement from satisfies . The expected counts are successes and failures, both at least . Therefore an approximately normal model for with mean and SD about is appropriate.”
If the condition fails, the unadjusted SD ignores important dependence; finite-population methods may be needed. If expected counts fail, do not use the usual normal approximation. In an independent setting, an exact binomial calculation or an appropriate simulation may still be available.
Quick check: does automatically pass the normality check when ?
No. and . The expected success count is below , so the usual normal approximation is not justified.
When is the shape nearly normal?
When is close to or , successes or failures are rare. A relatively large can still leave the sampling distribution skewed. Check both expected counts rather than using a blanket “” rule.
To make that visible, we generated independent samples at each of and , with success probability . Each sample contributes proportion. The random seed is 3,202,026.