Potential Errors When Performing Tests
A well-designed hypothesis test can still make a wrong decision. Learn to describe Type I and Type II errors, interpret power, and plan a study with the consequences of both errors in mind.
By the end of this lesson, you should be able to:
- Distinguish a test decision from the actual truth about a population.
- Identify Type I and Type II errors and describe them in context.
- Interpret , and power as conditional probabilities.
- Calculate and .
- Explain how sample size, variability, effect size and affect power.
- Connect the consequences of errors with choices made before collecting data.
Before you start: Review Topic 3.7’s complete hypothesis test. You should know , , -values, , and the decisions “reject ” and “fail to reject .”
First time learning this? Start with the decision table, then use the school example to put each error into words.
Here to revise? Use the error-and-power checklist, then try the practice before revealing answers.
The concept in 60 seconds
A school tests whether more than of its students prefer an earlier lunch. Let be the proportion of all students at this school who prefer the change.
A random sample supplies evidence, but it does not reveal the exact population proportion. Sometimes a sample looks unusually supportive even when the true proportion is . Sometimes it fails to show convincing support even when the true proportion is above .
Type I error: reject a true . The test finds evidence for an increase that is not actually there at the null benchmark.
Type II error: fail to reject when the specified alternative is actually true. The test misses a real increase.
These are errors in the statistical decision, not necessarily mistakes in arithmetic. You can write the hypotheses correctly, check the conditions and calculate the -value correctly, yet still make of these errors because samples vary.
Power describes how well the test can detect a specified real departure from . Higher power means a smaller chance of missing that departure.
All school, library and factory contexts are fictional teaching examples. Numerical planning illustrations use explicitly stated models or supplied probabilities; they are not measured performance claims.
Quick check: does every rejection of count as a Type I error?
No. It is a Type I error only when is actually true. Rejecting when the specified alternative is true is a correct decision. A decision alone does not tell us whether an error occurred.
Decision versus reality: possible outcomes
Keep two questions separate:
- What did the test decide? Reject or fail to reject .
- What is actually true? The null benchmark holds, or a specified departure in is real.
| Test decision | Reality: true | Reality: specified true |
|---|---|---|
| Reject | Type I error. Probability . | Correct detection. Probability . |
| Fail to reject | Correct nonrejection. Probability . | Type II error. Probability . |
Each probability is conditional on its column’s assumed truth. For the school’s right-tailed question, compare the null benchmark with a specified true increase such as . The displayed columns are distinct hypothetical population models; they are not categories whose probabilities add across a row.
The reality columns describe different possible populations. We do not usually know which column applies in an actual study. The table explains the possible outcomes; it does not reveal the hidden truth after test.
A memory aid that keeps the logic correct
- Type I: wrong rejection. We reject a null hypothesis that is true.
- Type II: missed departure. We fail to reject even though the alternative’s departure is real.
“False alarm” can help with a Type I error, and “missed detection” can help with a Type II error. But an error’s meaning depends on . If the alternative claims an improvement, the false alarm is a false claim of improvement; it is not automatically an alarm about a defect.
Failure to reject is not acceptance. It says that this test did not find convincing evidence for at the chosen level. It does not establish as true.
Quick check: the alternative is real, but the test rejects . Which error occurred?
No error. This is correct detection of the specified alternative. Its probability, under that true alternative, is the test’s power.
Describe both errors in context
A strong error description states what the test concludes or fails to detect, then states what the population is actually like. Avoid stopping at the labels “Type I” and “Type II.”
The test finds convincing evidence for when the school’s true is .
The test fails to find convincing evidence for when the school’s true is above .
Name the preference proportion and all students at this school. Failure to reject does not assert that .
The error description depends on . Reverse the direction for a less-than claim; include either direction for a not-equal claim.
School example: Type I error
“The test finds convincing statistical evidence that more than of all students at this school prefer an earlier lunch, when in fact the true preference proportion is .”
This is a wrong rejection of . The school could make a schedule change based on evidence that incorrectly suggests support exceeds the benchmark. The test does not guarantee any particular policy action; that is a separate decision.
School example: Type II error
“The test does not find convincing statistical evidence that more than of all students at this school prefer an earlier lunch, when in fact the true preference proportion is greater than .”
This is a missed real increase. The school could overlook a genuine preference for a change. Do not rewrite failure to reject as “the test concludes that at most prefer it.” That statement claims more than the test establishes.
For a less-than alternative
If a factory tests against , where is a batch’s defect proportion:
- Type I: find convincing evidence that the batch’s defect proportion is below when it is actually .
- Type II: fail to find convincing evidence of a reduction when the batch’s defect proportion is actually below .
For a not-equal alternative
If says , either an increase or a decrease can be a real departure. A Type I error claims a difference when . A Type II error misses a difference when really differs from .
Quick check: why is “Type II means concluding is true when it is false” an overstatement?
The test fails to reject ; it does not prove or conclude that is true. Describe Type II as missing convincing evidence for a specified real alternative.
, and power
The probabilities on this page are conditional: each assumes a particular truth about the population. The “given” condition is part of the meaning, not a detail to omit.
The vertical bar means “given” or “assuming.” Read each probability with its condition. is the significance level chosen before examining the data. , read “beta,” depends on the test design and the particular true alternative you want to detect.
Assume . The nominal probability of rejecting a true null is .
Assume . In the illustrated plan, about of samples fail to reject.
Assume that same . About of samples reject correctly.
These are the continuous school planning model’s probabilities explained in the next section. and power complement one another under the same true alternative. belongs to a different assumed truth.
| Term | Meaning | Illustrated school plan |
|---|---|---|
| : significance level | Nominal | under true |
| : Type II error probability | Approximately under true | |
| Approximately under the same true | ||
| Nominal probability of correct nonrejection under | under true | |
| -value | Observed extreme-result probability assuming the null model | Computed after observing a sample; not , or power |
Why do and power add to ?
Under the same specified true alternative, the test either rejects or fails to reject . Correct detection and a Type II error are complementary outcomes. For example, a supplied power of means .
Why do and not generally add to ?
They are calculated under different population truths. assumes is true. assumes a specified alternative is true. They are not the parts of probability total. A test might have and ; there is no requirement for their sum to equal .
Power is not the probability that is true
A power of at a specified true means that the test would reject in about of repetitions if that were the actual population proportion. It does not mean there is an chance that is true after observing data.
Different roles: is a planned threshold; a -value comes from the observed sample under the null model; power is a detection probability under a specified alternative model. None is the probability that a hypothesis itself is true.
In the AP test framework, represents the Type I error probability. For a discrete proportion test that uses a normal approximation, the actual rejection probability at the null benchmark can differ slightly from the nominal . The visual model below makes its approximation explicit.
Quick check: if power is at the stated true proportion, what is ?
. Under that same true proportion and design, there is a probability of failing to reject . Do not use to calculate power.
See the error probabilities
Consider a planned school study with an SRS of students from the school’s students, without replacement. The question remains versus , with selected in advance.
To examine detection, suppose the true population proportion were . That is an assumed alternative for this planning illustration. It is not a truth established by a sample.
The sampling checks are appropriate for the illustration: an SRS is specified, , and the expected counts are and under , or and under . Both distributions can be approximated by normal curves.