Estimators
A sample gives you a number. An estimator is the rule that produces it. Learn to estimate a population proportion, explain what “unbiased” means, and judge a method by how it behaves across repeated samples.
By the end of this lesson, you should be able to:
- Identify the population parameter a sample statistic estimates.
- Distinguish an estimator from the estimate obtained in sample.
- Calculate and interpret a sample proportion or sample mean.
- Use a sampling distribution’s mean to justify whether an estimator is unbiased.
- Separate bias, sampling variability and the error in estimate.
Before you start: Know population versus sample, proportions and sample means. Review Topic 2.12: Sampling Distributions and the Central Limit Theorem if you need help with repeated sampling.
First time learning this? Follow the school survey, then count all samples in the small-population example.
Here to revise? Use the quick checklist, then try the practice questions before opening the answers.
The concept in 60 seconds
A school wants to know the proportion of all its students who prefer an earlier lunch. Surveying everyone takes time, so it takes a random sample of students. Of those students, prefer an earlier lunch.
The sample proportion is . That single number is a point estimate of the unknown population proportion. It gives a useful starting answer, but a different sample would probably give a different number.
The key idea: An estimator is a method you can use on any sample. To judge whether it is unbiased, ask where its answers are centered across repeated samples—not whether answer happens to match the truth.
An unbiased estimator does not systematically overestimate or underestimate its target on average. Its sampling distribution has a mean equal to the population parameter it estimates.
All scenarios and numerical models in these notes are fictional teaching examples. Known population values are supplied so we can examine a method; in a real survey, the target value is usually unknown.
Quick check: does prove that exactly of all students prefer an earlier lunch?
No. It tells us that of the sampled students prefer it. We use that percentage to estimate the population percentage, while recognizing sampling variability and checking how students were selected.
Parameter, estimator and estimate
: the proportion of all school students who prefer an earlier lunch.
: count sampled students who prefer it, then divide by the sample size.
: the numerical answer for this sample.
A new sample can change and the estimate while the target population and estimation rule stay the same.
Parameter
A number describing a population. It is fixed for the population being studied, even if we do not know it.
: population proportion. : population mean.
Statistic
A number calculated from sample data. Its value can change from sample to sample.
, read “p-hat”: sample proportion. , read “x-bar”: sample mean.
Estimator
A statistic used as a rule for estimating a particular population parameter.
Use to estimate , where counts the sampled successes.
Estimate
The numerical value the estimator produces for the sample you actually observed.
With and , the estimate is .
The same notation often names both the estimator and its observed value. Context tells you whether we are discussing the rule before sampling or the number after sampling.
A point estimate is . A confidence interval adds a range to express uncertainty; that comes later in Unit 3. “Point” does not mean exact, certain or free from error.
Quick check: which is the estimator, or ?
is the rule. is the estimate from this particular school sample. Another sample uses the same rule but may produce a different estimate.
Calculate and interpret a point estimate
Estimating a population proportion
Choose category as a success. In statistics, “success” is just the outcome being counted; it does not have to be desirable. A defective item can be a success if the question concerns the proportion of defective items.
is the number of successes in the sample; is its total number of observations.
School example: .
Use the sample total as the denominator. Do not divide the sampled successes by the number of students in the entire school. That would mix a sample count with a population count.
A complete interpretation: “We estimate that of all students at this school prefer an earlier lunch, based on the random sample of students.” The target population and counted category both appear in the sentence.
Estimating a population mean
Here is the value of observation in the sample. Sample waits of , , and minutes give minutes.
The sample mean estimates the population mean wait. A mean keeps the original units; a proportion has no measurement units and can be written as a decimal or percentage.
Under a simple random sample from a finite population, the sample proportion is unbiased for the population proportion, and the sample mean is unbiased for the population mean. The same results hold for independent random observations from a common population model. The calculation alone does not guarantee a good sampling method.
Quick check: of inspected items are defective. What does “success” mean, and what is ?
A success is a defective item. . We estimate that of items in the target population are defective, assuming the inspection sample supports that population inference.
What makes an estimator unbiased?
Imagine taking repeated samples under the same sampling rule and sample size, calculating the estimator each time. Its sampling distribution describes the possible estimates and their probabilities.
Unbiased: the estimator’s sampling mean equals its target parameter.
Positive bias: the mean is above the target; the method overestimates on average.
Negative bias: the mean is below the target; the method underestimates on average.
You may see this written as for an unbiased estimator of a parameter . means the probability-weighted mean of ‘s sampling distribution. , read “theta,” is a general symbol for the target.
A complete example with just people
distinct people are labeled , , and . The Y labels mean “yes”; the N labels mean “no.” The true population proportion saying yes is .
Take a simple random sample of different people without replacement. There are equally likely unordered pairs. For each pair, calculate and an intentionally altered rule .
| Selected people | Probability of sample | ||
|---|---|---|---|
| , | |||
| , | |||
| , | |||
| , | |||
| , | |||
| , |
Each dot represents equally likely sample. Stacked dots preserve different samples giving the same result. The dashed line marks the true population proportion, , in both panels.
Average the values, including repeated values: . This equals , so is unbiased in this setting.
The altered rule has mean . Its mean is below , so it underestimates on average. Its bias is , or percentage points.
is an artificial rule chosen to illustrate bias. In general, under a sampling model with , its mean is and its bias is . A coincidence at one target value would not make a method unbiased for every population it is intended to estimate.
Quick check: samples give , but do not. Is still unbiased?
Yes. The criterion is the mean across the full sampling distribution, which is . An unbiased estimator does not need to equal the parameter for every sample.
When outcomes have unequal probabilities
In the example, every sample has probability , so averaging all results works. If you are given distinct estimator values with different probabilities, use a weighted mean instead.
| Value of | Probability | |
|---|---|---|