Justifying a Claim Based on a Confidence Interval for the Difference Between Two Population Means
You have a confidence interval. What can it tell you? Learn to interpret the difference, assess a claim and explain uncertainty without overstating what the data show.
By the end of this lesson, you should be able to:
- Interpret a confidence interval for a difference between population means in context.
- Use the interval’s position relative to to justify a claim.
- Explain what a confidence level means across repeated samples.
- Distinguish evidence of a difference from proof, causation and practical importance.
Before you start: Review the unpooled two-sample -interval in Topic 4.7. Know the difference between sample means and population means, and keep the subtraction order fixed.
First time learning this? Start with the delivery-time question, then compare positive, negative and -containing intervals.
Here to revise? Use the claim checklist, then attempt the practice questions before opening the solutions.
The concept in 60 seconds
A confidence interval estimates a difference between population means. The values inside the interval are compatible with the data under the chosen method and confidence level. Use the interval to judge whether the evidence supports the claim being asked about.
For the claim that the means differ: Check whether the interval contains . represents equal population means. If is outside a valid interval, it provides convincing evidence of a difference. If is inside, the interval does not provide convincing evidence that the means differ.
The signs explain the direction. For , an entirely positive interval supports A having the higher mean; an entirely negative interval supports B having the higher mean.
Containing does not prove equality. The interval may also contain meaningful positive and negative differences. You have uncertainty about the difference, not proof that it is exactly .
Does service A take longer on average?
Continue the delivery study from Topic 4.7. independently selected SRSs compare delivery times. The service-A sample has mean , and from deliveries. The service-B sample has mean , and from deliveries.
The independent samples, separate checks and both sample sizes being at least support the unpooled two-sample procedure. Population standard deviations are unknown. The confidence interval for is approximately .
Two questions, two jobs: “What does this interval estimate?” asks for an interpretation. “Is there convincing evidence that A has a longer mean delivery time?” asks for a claim justified with the interval.
Interpretation: We are confident that service A’s population mean delivery time is approximately longer than service B’s for the defined delivery populations.
Claim: The whole interval is above . Because every value in it represents a positive A-minus-B mean difference, it provides convincing evidence that A’s population mean delivery time is longer than B’s.
The conclusion concerns the population averages. It does not say that every A delivery takes longer than every B delivery.
Key ideas and notation
Parameter and order
is a fixed population mean difference. Define the populations, response and order before reading the signs.
For times in minutes, means A has the longer mean time.
Confidence interval
reports the estimated range of plausible values for using the chosen method. is the lower endpoint and is the upper endpoint.
Once calculated, endpoints are fixed numbers for that sample pair.
Claimed difference
A claim can concern , a direction such as , or a specific benchmark such as .
Translate the words into the parameter’s units and subtraction order.
Confidence level
, with the confidence level as a proportion, describes how often the method captures the true population mean difference over repeated random sampling under its conditions.
It is different from the endpoints of calculated interval.
Evidence and uncertainty
A conclusion should connect the interval to the claim using words such as “provides convincing evidence” or “does not provide convincing evidence.” An interval does not prove a population value, and values outside it are not declared impossible.
Write a clear interval interpretation
A useful interpretation has four pieces: the confidence level, the population means being compared, the endpoints and the measurement units.
We are [confidence level] confident that [defined population mean difference] is between and [units].
For a positive time interval, a sentence using “longer” can be easier to read. For a -containing interval, keep the signed difference or explain both directions.
Translate signs into everyday language
- for : A’s population mean time is estimated to be longer than B’s.
- for : A’s population mean time is estimated to be shorter than B’s.
- for : The estimate allows A’s mean to be up to shorter or up to longer than B’s, with equality also compatible.
Include the stated confidence level in a complete response. Do not quietly replace population means with sample means: the observed sample difference is already known.
Precision in wording: Say “the difference in population mean delivery times,” not just “the difference in delivery times.” The latter could sound like a claim about individual deliveries.
Read the signs and
For an interval estimating , is the no-difference benchmark. Read the entire interval, not only its midpoint.
| Interval position | What it supports | Example |
|---|---|---|
| Entirely above : | Evidence that . | |
| Entirely below : | Evidence that . | |
| Contains or reaches : | No convincing evidence of a difference from this interval. Equality is compatible, but unproven. |
These illustrative intervals come from different hypothetical comparisons, all expressed as in minutes. Dots mark interval midpoints. The dashed line is the equality benchmark.
What if a displayed endpoint is ?
If the unrounded interval reaches , do not describe it as entirely positive or negative. If an endpoint rounds to , check the unrounded value before deciding whether is actually included. A very small positive endpoint can look like after rounding.
Check yourself: A point estimate of is positive, but an interval crosses . The positive estimate alone does not settle the population’s direction.
Build a justified claim
State the populations, response and subtraction order. Translate “longer,” “lower” or “different” accordingly.
Compare , or another specified value, with endpoints. Check whether the whole interval supports the claimed direction.
Use “because the interval contains ” or “because the entire interval is above .”
State what the evidence supports about the population means. Keep uncertainty and the study’s scope clear.
This reasoning assumes the confidence interval comes from an appropriate procedure with justified conditions. Correct wording cannot repair an invalid interval.
Because [benchmark relationship], this interval [provides / does not provide] convincing evidence that [contextual claim about the population means].
Specific values and stronger claims
The delivery interval excludes , so it supports a positive mean difference. It also contains . Therefore, a difference of exactly is compatible with the interval; this does not prove the difference is .
The stronger claim “A’s mean is more than longer” requires checking whether the whole interval lies above . It does not: the lower endpoint is about . This interval alone does not establish that stronger threshold claim.
Optional: connection to a two-sided test
Let denote the confidence level as a proportion, with (so the confidence percentage is ). For the same unpooled procedure, data and degrees of freedom, a two-sided test of a specified mean difference at corresponds to a confidence interval with confidence proportion : a value strictly outside the unrounded interval has a matching two-sided -value below . At an exact endpoint, the matching -value equals ; the rule rejects at that boundary even though the endpoint belongs to the interval. Keep this boundary distinction separate from the classroom rule about whether the interval is entirely above or below .
For example, a interval excluding corresponds to rejection of a two-sided null of difference at . Do not automatically use this rule for a one-sided test at ; matching the tails and confidence level matters. Formal test setup comes in Topic 4.9.
Explain the confidence level
The population mean difference stays fixed for the defined populations. The random samples change, so their means, standard deviations and confidence interval endpoints change.
confidence means: If we repeatedly took independent random samples of the same sizes from the same populations and used this interval method, approximately of the intervals would capture the true difference between the population means, under the procedure’s conditions.
For calculated interval, the fixed population difference is either inside or outside. We usually do not know which. Confidence describes the reliability of the method, not a probability assigned to the fixed parameter after the interval has been calculated.