AP Statistics / Unit 3: Inference for Categorical Data: Proportions / Topic 3.2
NUM8ERS study notes · Topic 3.2

Sampling Distributions for Sample Proportions

Different random samples give different proportions. Learn where those proportions are centered, how much they vary, when a normal curve is appropriate, and how to calculate and explain a probability about a sample.

2026–27 curriculum6 worked examples10 practice questionsNormal curves + sampling visuals

By the end of this lesson, you should be able to:

  • Distinguish a population proportion pp from a sample proportion p^\hat{p}.
  • Calculate the mean and standard deviation of p^\hat{p} under a suitable sampling model.
  • Check random sampling, the 10%10\% condition and expected success/failure counts.
  • Describe how sample size affects the spread and shape of p^\hat p‘s sampling distribution.
  • Calculate normal-model probabilities and interpret them in the sampling context.

Before you start: Know proportions, sampling distributions and normal probabilities. Review Topic 3.1: Estimators for unbiasedness and Topic 2.11: The Normal Distribution for z-scores and areas.

First time learning this? Follow the school survey, check its conditions, then work through the shaded probability example.

Here to revise? Use the formula checklist, then try the practice questions before revealing solutions.

The concept in 60 seconds

Imagine a school of 10,00010{,}000 students where 60%60\% prefer an earlier lunch. The population proportion is p=0.60p = 0.60. Take a simple random sample of 100100 students without replacement. If 6464 say yes, the sample proportion is p^=64100=0.64\hat{p} = \frac{64}{100} = 0.64.

A fresh random sample of 100100 might give 0.57,0.610.57, 0.61 or another value. The sampling distribution of p^\hat{p} describes all possible sample proportions and how likely they are under the specified sampling method and sample size.

Visual guide 1: responses → 11 statistic → repeated-sample distribution
11 sample100100 individual responses

Each student gives 11 yes/no response.

Suppose 6464 say yes.

11 summary11 sample proportion

Calculate p^=64100=0.64\hat{p} = \frac{64}{100} = 0.64.

This sample contributes 11 plotted value.

Repeat the processMany sample proportions

Take fresh samples of the same size and record 11 p^\hat{p} each time.

The graph describes how those answers vary.

Keep the population, sampling rule and nn fixed while constructing one sampling distribution. A histogram of 11 sample’s raw binary responses is a different graph.

Remember the target: pp stays fixed for this population. p^\hat{p} varies from sample to sample. A graph of repeated p^\hat{p} values is a graph of sample summaries, not a graph of individual students’ yes/no responses.

The proportions tend to cluster around pp. Larger random samples usually give a tighter cluster. With enough expected successes and failures, we can use an approximately normal sampling model.

All scenarios here are fictional teaching examples. Population proportions are supplied so we can study sampling behavior. The shape comparison uses an actual seeded simulation from a specified model; it is not collected school data.

Quick check: a sample of 100100 students gives how many p^\hat{p} values?

11. Calculate the number saying yes divided by 100100. To plot 5,0005{,}000 sample proportions, repeat the sampling process 5,0005{,}000 times and record 11 proportion each time.

Notation and the sampling model

The school’s population and 11 observed sample.
SymbolMeaningSchool example
ppPopulation proportion in the counted category0.600.60
p^\hat{p}Proportion in 11 sample; Xn\frac Xn0.640.64
XXSuccess count in 11 sample6464 students
nnNumber of observations per sample100100 students
NNSize of the finite population10,00010{,}000 students

A “success” is the category being counted. It might mean preferring an earlier lunch, owning a bicycle, or receiving a defective item. The word does not imply that the outcome is desirable.

p^=Xn\hat p=\frac Xn

School sample: p^=64100=0.64\hat p=\frac{64}{100}=0.64.
School population: p=6,00010,000=0.60p=\frac{6{,}000}{10{,}000}=0.60.

Connect a count to a proportion

For nn independent binary observations with the same success probability pp, the success count is X∼Binomial⁡(n,p)X \sim \operatorname{Binomial}(n, p), and p^=Xn\hat{p} = \frac Xn. The binomial count and the sample proportion describe the same sample on different scales.

For n=100n = 100, possible proportions are 0,0.01,0.02,…,10,0.01,0.02,\ldots,1. The true sampling distribution is discrete. A normal curve is a continuous approximation to those possible values.

In a finite population sampled without replacement, the observations are dependent. A small sampling fraction lets us approximate the process with independent draws. If the fraction is substantial, use a model that accounts for sampling without replacement; the binomial model is then not exact.

Quick check: with n=100n = 100, what count corresponds to p^≥0.70\hat{p} \ge 0.70?

X≥70X \ge 70. The event p^>0.70\hat{p} > 0.70 instead corresponds to X≥71X \ge 71. Including a boundary matters for the exact discrete distribution, even though a continuous normal curve assigns a single point 00 area.

Center, spread and sample size

For independent observations with the same population success probability pp:

Under the independent model:

μ(p^)=p\mu(\hat p)=p
σ(p^)=p(1−p)n\sigma(\hat p)=\sqrt{\frac{p(1-p)}{n}}

The mean equals pp, so p^\hat{p} is an unbiased estimator under this model. The SD measures how much sample proportions vary across repeated samples. These formulas do not require a normal shape; normality is a separate question.

Apply the formulas to the school

Here p=0.60p = 0.60 and n=100n = 100. The sampling fraction is small enough to support the independent approximation, as we will check below.

Mean is 0.600.60. The school’s sampling SD is approximately

σ(p^)≈0.60(0.40)100=0.0024≈0.04899\sigma(\hat p)\approx\sqrt{\frac{0.60(0.40)}{100}}=\sqrt{0.0024}\approx0.04899

The model is centered at 60%60\%, with a sampling SD of about 4.904.90 percentage points. It is not centered at the observed p^=0.64\hat{p} = 0.64. We use the supplied pp to describe the model for repeated samples.

What changes when nn increases?

Visual guide 2: a larger sample narrows the proportion’s distribution
Sample size and the sampling spread of a proportionTwo approximate normal density curves share axes and a center of 0.60. The n = 100 curve has SD about 0.049. The n = 400 curve has SD about 0.0245 and is narrower and taller. A dotted line marks the common center. 0.4 0.5 0.6 0.7 0.8 Sample proportion p̂ 0 5 10 15 20 Density Same center; smaller spread · normal approximations n = 100; SD ≈ 0.049 n = 400; SD ≈ 0.0245 Sample size and the sampling spread of a proportionTwo approximate normal density curves share axes and a center of 0.60. The n = 100 curve has SD about 0.049. The n = 400 curve has SD about 0.0245 and is narrower and taller. A dotted line marks the common center. 0.4 0.5 0.6 0.7 0.8 Sample proportion p̂ 0 5 10 15 20 Density Same center; smaller spread Normal approximations n = 100; SD ≈ 0.049 n = 400; SD ≈ 0.0245

Both are theoretical normal approximations for the same p=0.60p = 0.60; they are not simulation histograms. The shared horizontal range is 0.35–0.850.35\text{–}0.85. Density height is not a probability; probability is area under a curve. The full normal curves each have total area 11.

For n=400n = 400, the approximate SD is 0.24400≈0.02449\sqrt{\frac{0.24}{400}} \approx 0.02449, or 2.452.45 percentage points. The center remains 0.600.60. Multiplying nn by 44 halves the SD because nn is inside a square root.

Increasing sample size changes the statistic’s sampling distribution. Increasing the number of simulation repetitions at fixed nn only gives a more stable picture of that same distribution.

Do not swap pp for p^\hat{p} automatically: This lesson describes a sampling distribution using a supplied population proportion pp. Estimating its SD from sample data uses a standard error involving p^\hat{p}; that is developed with confidence intervals in Topic 3.3.

Quick check: if nn doubles, does the SD halve?

No. It is multiplied by 12≈0.707\frac{1}{\sqrt{2}}\approx0.707. To halve the SD while keeping pp fixed under the same independent model, multiply nn by 44.

Check the conditions

Visual guide 3: three checks, three different jobs
SelectionRandom sample

Ask: How were the observations selected?

A simple random sample supports inference to the intended population.

Dependence10%10\% condition

When sampling without replacement: n≤0.10Nn \le 0.10N.

A small sampling fraction supports the usual independent-model SD approximation.

ShapeExpected counts

Check both: np≥10np \ge 10 and n(1−p)≥10n(1 – p) \ge 10.

This supports an approximately normal sampling distribution.

Passing one check does not establish the others. Use pp from the supplied population model for the expected counts in this lesson.

1. Random sampling

The data should come from an appropriate random sample of the target population. In the school example, the sample is a simple random sample. A large volunteer sample would not satisfy this selection requirement simply because it contains many people.

2. The 10%10\% condition when sampling without replacement

n≤0.10N⟺N≥10nn\le0.10N\quad\Longleftrightarrow\quad N\ge10n

School check: 100≤0.10(10,000)=1,000100\le0.10(10{,}000)=1{,}000. The condition is met.

This condition supports treating the dependence from sampling without replacement as small when using p(1−p)n\sqrt{\frac{p(1 – p)}{n}}. It does not make the observations exactly independent, and it does not establish a normal shape.

For genuinely independent sampling, such as a specified independent binary-response model or sampling with replacement, this finite-population check is not needed. Independence must still be justified by the process.

3. Enough expected successes and failures

np≥10andn(1−p)≥10np\ge10\quad\text{and}\quad n(1-p)\ge10

School check: 100(0.60)=60100(0.60)=60 expected successes; 100(0.40)=40100(0.40)=40 expected failures. Both are at least 1010.

These counts support an approximately normal shape. Use the supplied population proportion pp. The observed sample may contain 6464 successes, but the expected count under p=0.60p = 0.60 is 6060.

A full justification for the school: “The students form a simple random sample. Sampling 100100 without replacement from 10,00010{,}000 satisfies 100≤1,000100 \le 1{,}000. The expected counts are 6060 successes and 4040 failures, both at least 1010. Therefore an approximately normal model for p^\hat{p} with mean 0.600.60 and SD about 0.048990.04899 is appropriate.”

If the 10%10\% condition fails, the unadjusted SD ignores important dependence; finite-population methods may be needed. If expected counts fail, do not use the usual normal approximation. In an independent setting, an exact binomial calculation or an appropriate simulation may still be available.

Quick check: does n=100n = 100 automatically pass the normality check when p=0.02p = 0.02?

No. np=2np = 2 and n(1−p)=98n(1 – p) = 98. The expected success count is below 1010, so the usual normal approximation is not justified.

When is the shape nearly normal?

When pp is close to 00 or 11, successes or failures are rare. A relatively large nn can still leave the sampling distribution skewed. Check both expected counts rather than using a blanket “n≥30n \ge 30” rule.

To make that visible, we generated 5,0005{,}000 independent samples at each of n=20n = 20 and n=100n = 100, with success probability p=0.10p = 0.10. Each sample contributes 11 proportion. The random seed is 3,202,026.

Visual guide 4: rare successes need a larger sample
Shape of two simulated sample-proportion distributionsTwo histograms from 5,000 independent samples each with p = 0.10. At n = 20, expected successes are 2 and the normality check fails; the sample proportions are discrete and right-skewed. At n = 100, expected successes are 10, the check passes, and the distribution is more concentrated with some remaining right skew. 0.0 0.1 0.2 0.3 0.4 Sample proportion p̂ 0% 20% 40% Share of 5,000 n = 20; expected yes = 2 · normality check fails 0.0 0.1 0.2
Posted on Google Google
0000003998 : Abdad Alam Shamim Alam profile picture
0000003998 : Abdad Alam Shamim Alam
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
This is the bestest institution I have ever come to and I love it very much.
Posted on Google Google
Aliki S profile picture
Aliki S
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Ali profile picture
Ali
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Sly is him
Posted on Google Google
Jitendra Kumar Kumawat profile picture
Jitendra Kumar Kumawat
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Raahil Hasan profile picture
Raahil Hasan
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Tina Mudarres profile picture
Tina Mudarres
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
The classes are amazing and my child learnt so much
Posted on Google Google
alisha gadoya profile picture
alisha gadoya
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I have experienced a lot of good things, it has taught me so many things and from B/Cs I have gone to an A! This is wonderful and Mavish’s class is awesome.
Posted on Google Google
smasher 123 profile picture
smasher 123
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
It’s sensational my kid went. First he only would get C now it’s all A’s really good and recommend
Posted on Google Google
Mosa Al- Samaraie profile picture
Mosa Al- Samaraie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
AMAZING PRICES GREAT TEACHERS STRAIGHT FORWARD LEARNING STEADY PACE IN TUTORING
Posted on Google Google
Cael Dagnelie profile picture
Cael Dagnelie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I had a great experience learning math with Ms. Mavish. She explains complex topics in a very clear and simple way, which made it easier for me to understand and enjoy the subject. Her patience and dedication really stood out, and she always made sure that everyone in the class was keeping up. I especially appreciated how approachable she was — I never felt afraid to ask questions, and she was always willing to help. Thanks to her teaching, my confidence in math has grown a lot. I’m really thankful for the effort she puts into every lesson!

NUM8ERS is one of finest tutoring institutes in UAE, Located in Al Barsha 1, Dubai. Close to DUBAI AMERICAN ACADEMY (DAA) & AMERICAN SCHOOL OF DUBAI (ASD).