Setting Up a Test for the Difference Between Two Population Means
Does group have a higher average, a lower average, or simply a different average? Turn that question into a clear two-sample -test setup before calculating anything.
By the end of this lesson, you should be able to:
- Choose a two-sample -test for independent groups with a quantitative response.
- Define both population means and write the null and alternative hypotheses.
- Choose the alternative’s direction from the original question.
- Justify randomization, independence and the sample-size or shape conditions.
Before you start: Review population versus sample means, independent versus paired data, and Topic 4.8: using a confidence interval to justify a claim.
First time learning this? Follow the delivery example, then build the four complete test setups.
Here to revise? Use the setup checklist and attempt the practice questions before revealing solutions.
The concept in 60 seconds
A two-sample -test asks whether data from independent groups provide convincing evidence of a difference between their population means. The response must be quantitative, such as time, height or score.
The setup has three jobs: Name the population means, state the competing claims about their difference, and check whether the study supports the test.
The null hypothesis is the starting benchmark: no difference in the population means. The alternative states the difference you are looking for: greater than, less than, or different from .
This lesson builds that plan. A sample mean difference alone does not tell you whether the population means differ. The test statistic, -value and final decision come in Topic 4.10.
Does service A take longer on average?
A delivery company asks whether service A’s population mean delivery time is longer than service B’s. It independently selects an SRS of of service A’s deliveries and an SRS of of service B’s deliveries from a defined period.
| Service | Sample size | Sample mean | Sample |
|---|---|---|---|
| A | |||
| B |
Let and be the population mean delivery times, in minutes, for these services during that period. Keep the order .
The alternative is positive because the original question asks whether A takes longer. The observed difference, , is a sample result; it is not the value to put in the alternative hypothesis.
The samples are independently selected, each is at most of its delivery population, and both sample sizes are at least . These facts support an independent two-sample -test. We have a valid setup, but have not yet calculated the evidence against the null.
Key ideas and notation
Population means: and
The unknown average responses for the defined populations. Specify who or what belongs to each population, the response and its units.
is the population mean time for service A’s deliveries during the study period.
Mean difference:
is the population parameter being tested. Positive and negative values depend on which group comes first.
means equal population means; it does not mean identical individuals or distributions.
Sample means: and
These are calculated from the observed data. Their difference estimates and later helps measure evidence against .
Hypotheses concern and , not the already observed and .
Null and alternative
states population mean difference. states the direction or difference the investigation seeks.
Choose before inspecting the sample outcome.
Independent two-sample -test
A test for a difference between population means using independent groups and sample standard deviations because population standard deviations are unknown. The unpooled method allows the population variances to differ.
Choose the method before the hypotheses
Start with the response and how the observations were collected. Having columns of numbers does not automatically mean you have independent samples.
Compare population mean with a stated value. Example: mean package weight versus .
The same people measured times, or deliberately matched pairs. Compute difference per pair.
Separate groups without within-pair matching. Compare their population mean quantitative responses.
Success/failure or yes/no outcomes concern proportions. A test for means does not fit that response.
After choosing a method, check its conditions. This diagram identifies the design; it does not guarantee that inference is valid.
Test or interval?
Use a confidence interval to estimate a population mean difference. Use a hypothesis test to assess evidence against a specified claim. “How much higher?” calls for estimation; “Is there convincing evidence it is higher?” calls for testing.
Why , and why unpooled?
In these comparisons, the population standard deviations are unknown and the sample standard deviations estimate them. The procedure accounts for that estimation. The usual unpooled two-sample method does not assume equal population variances. Equal-looking sample standard deviations are not a reason to silently switch to a pooled procedure.
Write the null and alternative hypotheses
First define the parameters in a sentence. Then write both hypotheses using the same subtraction order.
Equivalent form:
The null includes equality. Choose one alternative from the original question:
| Investigative question | Alternative | Type |
|---|---|---|
| Is group 1’s population mean higher? | One-sided | |
| Is group 1’s population mean lower? | One-sided | |
| Are the population means different? | Two-sided |
A parameter definition that earns its place
“ is group 1” is incomplete. Write: “ is the population mean task-completion time, in seconds, for users of the new interface.” Define just as precisely for the old interface.
Check the response: For time, “faster” usually means a lower mean. For score, “better performance” may mean a higher mean. Translate the meaning of the measurement, not just a positive-sounding word.
Do not write merely because the observed sample difference is . A directional claim covers a range of possible population differences, and the population value remains unknown.
Choose the direction from the question
A one-sided alternative looks for a difference in specified direction. A two-sided alternative looks for a difference in either direction. The study question decides this choice before the results are examined.
. These highlighted number-line regions show parameter values allowed by each alternative. They are not -values, probability distributions or rejection regions. The open circle excludes .
Reversing the subtraction order
If the claim is that A’s mean time is longer, you may write or . Both describe the same claim. Define your order and keep it consistent throughout the test.
Reversing the order changes the signs, including the sign of the later test statistic. It does not change the evidence for the same substantive claim when the alternative is reversed consistently.
Think before revealing: A study asks whether brands have different mean battery lives. The first sample mean happens to be larger. Should the alternative become “greater than”?
Reveal the answer
No. “Different” calls for a two-sided alternative, . Changing to a one-sided test after seeing the result changes the question to favor the observed direction.
Check the conditions separately
Do more than name a checklist. Connect each condition to a fact about the actual study. Check the groups separately where appropriate.
1. Randomization and independent groups
Use independently selected random samples or a properly conducted randomized experiment assigning individuals to independent treatment groups. Repeated measurements on the same people are paired and require a different analysis.
Random sampling supports generalization to the sampled populations. Random assignment supports a causal comparison of treatments. The two serve different purposes. A large convenience sample does not become random simply because it has many observations.
2. The condition when sampling without replacement
and
Each sample should be at most of its own population to support treating the sampled observations as approximately independent within that group. For the delivery study: and .
This sampling-without-replacement check is not required merely because a randomized experiment assigns participants to treatments. Do not invent a population size for a random-assignment problem. If a study also uses random sampling, assess that sampling step where relevant.
3. Sample size or distribution shape
The sample-size condition is met if both and . Alternatively, both populations may be known to be approximately normal.
If either sample has fewer than observations and population normality is not established, inspect both sample distributions: neither should show strong skewness or outliers. A small sample needs convincing shape evidence; do not claim that an unknown population is normal merely because is small.