AP Statistics / Unit 1: Exploring One-Variable Data and Collecting Data / Topic 1.12
NUM8ERS study notes · Topic 1.12

Potential Problems with Sampling

A survey can collect lots of answers and still give a misleading picture. Learn to spot who is missing, who chooses to take part, and how the questions can push the answers in one direction.

2026–27 curriculum4 worked examples8 practice questionsBias + visual guides

By the end of this lesson, you should be able to:

  • Explain the difference between chance variation and systematic bias.
  • Recognize undercoverage, voluntary response, nonresponse and response bias.
  • Identify why convenience sampling can miss important parts of a population.
  • Explain a possible overestimate or underestimate using the study’s details.
  • Suggest an improvement that addresses the actual source of the problem.

Before you start: Know the population, sampling frame and selected sample. Review Topic 1.11: Random Sampling if those ideas need a refresh.

First time learning this? Follow the school survey through its stages, then compare the four sources of bias.

Here to revise? Use the diagnosis checklist, then try the practice questions before opening their solutions.

The concept in 60 seconds

Bias is a systematic tendency for an estimate to miss its target in a particular direction. A procedure might tend to make an average too low or a proportion too high. The issue is how the data are obtained, not simply how many observations there are.

Ask three practical questions: Could everyone of interest be selected? Who actually supplied data? Are the recorded answers accurate? Those questions help locate a problem before you calculate a mean or proportion.

Follow the survey from the population to the reported result
1PopulationWho do we want to learn about?
2FrameWho is available for selection?
3Selected sampleWho is chosen, and by what rule?
4RespondentsWho actually provides data?
5Recorded answersWhat do the responses really measure?

A good selection rule is one step. Missing people, missing responses or distorted measurements can still affect the result.

A larger sample does not repair a flawed process. Asking more people from an incomplete frame does not bring the missing group onto that frame. Collecting more answers to a leading question does not make the wording neutral.

A potential source of bias does not tell you its exact size. You may be able to explain a likely direction, or you may need more information. A careful answer identifies the mechanism and avoids inventing what unobserved people would have said.

Cedar High reviews its travel-time survey

Cedar High wants to estimate the mean one-way home-to-school travel time, in minutes, for all 1,2001{,}200 enrolled students on a specified Tuesday. Students travel in different ways and live at different distances. Each proposal below targets that same population and measurement.

Four fictional proposals. Find the weak point in each collection process.
ProposalWhat happens?The concern to investigate
Use the parking-permit listSelect 120120 students randomly from permit holders, then describe the result as a school-wide estimate.Students without permits cannot be selected from this frame. Their travel times may differ.
Post an open survey linkAny enrolled student may click a link headed “Tell us about frustrating journeys.” Only students who opt in supply data.Interest in the topic may affect who volunteers. This is not a random sample of enrolled students.
Stop after 1 invitation1\text{ invitation}Take an SRS of 120120 from a complete enrollment list. Only 4242 answer the first invitation; the other 7878 are ignored.The responding students may differ from the selected students who did not answer.
Push a short-time answerAsk selected students, “Your journey usually takes under 20 min20 \,\mathrm{min}, doesn’t it?” and use the answers to describe Tuesday travel.The question suggests a preferred answer and changes a specific Tuesday measurement into “usually.”

Randomizing the selection from a parking list does not solve its missing coverage. Randomizing from the full list does not ensure everyone responds. Even with full response, misleading wording can distort the recorded data.

Think first: Which proposal starts with a good random selection process but loses some of the selected students before obtaining data?

Check your prediction

The third proposal. The original 120120 are an SRS from the enrollment list, but data come from only 4242 respondents. The main concern is possible nonresponse bias, especially if the 7878 nonrespondents have systematically different Tuesday travel times. Their silence is not a travel time of 0\text{travel time of }0.

Bias versus chance variation

Two well-conducted random samples can produce different means even when neither method systematically favors high or low values. That sample-to-sample movement is chance variation. Bias describes a repeated tendency of a procedure, rather than the accidental difference in one sample.

Different questions about an estimate
Chance variationDoes the estimate move around?
Below target ↔ Above target

Different random selections give different results. In a suitable unbiased procedure, there is no persistent shift away from the target.

Systematic biasIs the process tilted?
Process → tendency too high or too low

A missing group or a distorted answer can repeatedly pull the estimate away from the population value.

This is a conceptual comparison, not a plot of simulated results. A biased procedure can still produce an occasional estimate near the true value.

The parameter is the population value of interest, such as the true mean Tuesday travel time. A statistic is calculated from the collected data, such as the respondents’ mean. You usually do not know the parameter, so you cannot measure the exact error just by inspecting the statistic.

One estimate is too highThis alone does not show that the sampling method is biased. A random sample can be above the target by chance.
The method omits a relevant groupThis supplies a reason to suspect bias, even before the true population mean is known.

Increasing the size of a properly collected random sample generally reduces random sampling variability. It does not automatically remove a systematic problem. A large online poll may estimate the opinions of its volunteers very precisely while giving a poor estimate for the entire population.

The vocabulary here concerns statistical bias. A collector does not have to intend to mislead: a convenient choice, confusing wording or an outdated list can create a problem accidentally.

Recognize the main sources of bias

Undercoverage: part of the population is poorly represented

Undercoverage is a concern when a selection process leaves out part of the target population or makes a relevant part less likely to be selected. A list of parking-permit holders is an incomplete frame for all enrolled students. Bus riders, walkers and other students without permits cannot enter a sample selected only from that list.

Explain why the missing group matters for the variable being estimated. If students without permits tend to have longer travel times, a permit-only estimate may be too low. If you know nothing about their times, identify the coverage problem without claiming a direction.

Improve the process: use a current enrollment frame covering all students and an appropriate random selection rule. Increasing the number chosen from the permit list leaves the same group excluded.

Voluntary response: people choose themselves into the sample

An open invitation followed by a sample made entirely of people who opt in creates a voluntary response sample. People with a strong interest, strong feelings or more time may be more likely to take part. A journey-complaints poll may especially attract students unhappy with their travel.

The possibility of bias comes from self-selection and its connection to the response. Volunteers are not automatically dishonest. 100 honest answers100\text{ honest answers} from a particular kind of volunteer can still give an unbalanced picture of a much larger group.

Improve the process: select units using chance from the target population’s frame, then contact those units and use a thoughtful follow-up plan. Do not treat the first people willing to answer as a substitute for that selection process.

Nonresponse: selected units fail to provide data

Nonresponse occurs when some chosen units do not provide the requested data. They may not be reached, may decline, or may leave a relevant item blank. Nonresponse bias is a concern when respondents and nonrespondents differ in a way related to the study’s variable.

For Cedar High, suppose students with long journeys are especially unlikely to answer because the survey arrives during their commute. Their absence could pull the respondents’ mean below the full population mean. A low response rate raises concern, but the rate alone does not prove the direction or amount of bias.

Improve the process: follow up with the originally selected students, offer suitable response times or modes, and make the survey easy to complete. Replacing unavailable students with nearby volunteers changes the responding group and does not automatically fix the design.

Response bias: the supplied values are systematically distorted

A person can be selected and respond, yet supply an answer that differs from the true value. Leading questions, unclear wording, memory problems or pressure to give a socially acceptable answer can create a systematic shift. Measurement procedures can also be at fault: a scale that reads every object’s mass 2 g2 \,\mathrm{g} too high creates a consistent upward error.

Ask whether the answers tend to be pushed in a particular direction. Ordinary mistakes can occur in either direction; not every isolated error establishes response bias. In a survey about arriving late, a student may underreport delays if answering directly to the staff member who records lateness.

Improve the process: define the measurement clearly, use neutral questions, choose a suitable recall period, provide a response setting that reduces pressure, and check instruments. Do not claim that anonymity or any single change guarantees perfectly accurate responses.

Similar-looking problems happen at different stages
UndercoveragePopulation → incomplete frame

Some relevant students cannot enter the selection pool.

Ask: Were they available for selection?

Voluntary responseOpen invitation → self-selected volunteers

People choose themselves into the data set.

Ask: Who decided which people would participate?

NonresponseChosen units → some missing answers

Some selected students do not supply the data.

Ask: Were they chosen but did not answer?

Response biasRespondent → distorted recorded value

An answer is present, but may be systematically inaccurate.

Ask: Did the response measure the truth?

Convenience sampling: choose whoever is easy to reach

Asking the first 6060 students near 11 gate is a convenience sample. Ease of access determines inclusion. Students using another entrance, arriving at a different time or absent that morning may be missed. If arrival patterns are related to travel time, the result can be systematically unrepresentative.

Convenience and voluntary response are both nonrandom selection methods, but the mechanisms differ: the collector chooses accessible units in the first; participants opt into an open invitation in the second. A study can contain both mechanisms, along with other problems.

Giving consent is not the same as creating a voluntary response sample. In an SRS survey, people are selected first and remain free to decline. Missing answers then create a nonresponse concern. In an open poll, people choose themselves into the sample in the first place.

Not every unequal selection chance signals a bad design. Topic 1.11 allowed different sampling fractions in strata, with suitable analysis. Here the concern is a selection or collection process that systematically misrepresents the target, rather than any planned unequal probability by itself.

Improve questions and measurement

A question should make the intended variable clear without suggesting the answer the researcher wants. For a travel-time measurement, specify the journey, day, starting point, ending point and units. “Tuesday” and “a typical week” are different targets.

From a pushed answer to a defined measurement
Problematic wording
“Your journey usually takes under 20 min20 \,\mathrm{min}, doesn’t it?”

Suggests a short journey, asks for a yes/no judgment, and uses an undefined “usually.” It does not collect the specific Tuesday duration.

Improved wording
“On Tuesday 6 October, how many minutes passed from leaving home to arriving at school for your morning journey?”

Names the day, journey and units without proposing a preferred duration. In a real survey, replace the example date with the actual study date.

These are original teaching examples. Clear wording helps; recall and recording errors may still occur.

Small wording choices can change what an answer represents.
ProblemExampleA better approach
Loaded description“Do you support the wasteful new bus route?”Describe the route neutrally, then ask for support, opposition or no opinion.
Two issues in one item“Are the buses punctual and comfortable?”Use separate punctuality and comfort questions. A respondent could answer differently to each.
Unclear reference period“How often are you late?”Specify a period and define what counts as late, such as arriving after the scheduled start on each school day last week.
Unbalanced response choicesOnly “good,” “very good” and “excellent” are offered for a service rating.Allow ratings below and above neutral, with applicable options such as “did not use the service.”

Self-reported data are not automatically useless. Ask whether people can recall the information and whether the setting encourages accurate reporting. Recording a journey’s start and finish when it happens may reduce recall problems compared with guessing its duration weeks later.

A neutral question does not restore missing population coverage. A complete frame does not calibrate a faulty instrument. Match the improvement to the stage where the error occurs, and keep that distinction in your explanation.

Four worked examples

All scenarios and quantities below are invented for teaching. The reasoning shows how a collection process could affect an estimate; it does not report findings from real surveys.

Example 1: Undercoverage can shift a mean

Task: a fictional school has 900900 nearby students with a population mean Tuesday travel time of 20 min20 \,\mathrm{min} and 300300 farther-away students with a population mean of 40 min40 \,\mathrm{min}. A researcher randomly samples only from the nearby-student list, but wants to estimate the mean for all 1,2001{,}200 students.

  1. Name the target: the mean travel time of all 1,2001{,}200 enrolled students.
  2. Locate the missing group: the 300300 farther-away students are absent from the frame, so they cannot be selected.
  3. Use the supplied group information: farther-away students have the higher mean. Their exclusion removes a group that should pull the school-wide mean upward.
  4. State the possible effect: a nearby-only estimate tends to underestimate the full school’s mean.
Whole-school mean=900(20)+300(40)1,200=25 min\text{Whole-school mean}=\frac{900(20)+300(40)}{1{,}200}=25\,\mathrm{min}

Covered-group mean=20 min\text{Covered-group mean}=20\,\mathrm{min}; difference from the target=20−25=−5 min\text{difference from the target}=20-25=-5\,\mathrm{min}.

The group means are known only because this is a constructed example. Random sample means from the nearby group may vary around 20 min20 \,\mathrm{min}, but collecting more nearby students does not include the farther-away group. The 55-minute gap shows the coverage problem’s direction in this illustration; it is not a formula for calculating unknown bias in an ordinary survey.

Improve: select randomly from the complete enrollment list, or use an appropriate design covering both groups. If stratifying, use the population shares correctly in the whole-school estimate.

Example 2: Nonresponse is about chosen people

Task: a library selects an SRS of 200200 registered members to estimate mean books borrowed during September. 80 members80\text{ members} respond. The study description states that frequent borrowers are more likely to respond than infrequent borrowers. Identify the concern and its likely direction.

  1. Separate selection and response: 200200 members were selected; only 8080 supplied data.
  2. Name the source: possible nonresponse bias. The concern arises after the random selection, because selected members do not all answer.
  3. Connect the missing answers to the variable: the responding group disproportionately contains frequent borrowers.
  4. State the direction: the respondents’ mean number borrowed may overestimate the mean for all registered members.

The response fraction is 80200=40%\frac{80}{200}=40\%, and 120120 selected members did not respond. These counts alone would not establish an upward bias. The stated difference in borrowing behavior supports that direction.

Improve: follow up with the originally selected members using suitable contact methods. Where appropriate and permitted, a consistent borrowing-record measurement might reduce reliance on remembered counts. Do not record unanswered surveys as 0 books borrowed0\text{ books borrowed}.

Example 3: Voluntary responses do not represent everyone

Task: a recreation center posts “Unhappy with our booking system? Tell us!” on its website. Of the 500500 visitors who submit the optional poll, 70%70\% report dissatisfaction. The manager concludes that 70%70\% of all registered members are dissatisfied.

  1. Name the target population: all registered members, not just website visitors who answered.
  2. Identify how inclusion happened: people chose themselves into an open poll. This creates a voluntary response sample.
  3. Use the invitation’s focus: it specifically attracts people unhappy with booking. Dissatisfied members may be more motivated to participate.
  4. Assess the conclusion: 70%70\% describes the poll’s respondents. Applying it to all members may overestimate the dissatisfied proportion.

There may also be undercoverage if some registered members do not use the website. The leading invitation adds another concern. When several problems occur together, identify the requested source and support it with the actual selection or wording detail.

Improve: take a random sample from a complete member frame, contact the selected members, use a neutral invitation and question, and follow up thoughtfully. 5,000 self-selected responses5{,}000\text{ self-selected responses} to the same poll would still leave its selection problem.

Example 4: An answer can be present but inaccurate

Task: students are randomly selected from a complete list. Every selected student answers a survey about days late to school last week, face-to-face with the staff member who issues lateness notices. The prompt states that students tend to report fewer late days to avoid embarrassment.

  1. Check for missing people: the frame is complete and all selected students respond; the stated problem is not nonresponse.
  2. Check the recorded values: students systematically underreport their actual number of late days.
  3. Name the source: response bias, associated here with pressure to give a socially acceptable answer.
  4. State the direction: the mean reported number of late days may underestimate the true mean for the target students.

Improve: define the period clearly and use a suitable confidential collection setting that reduces this pressure. A consistently defined attendance-record measure may be useful if appropriate for the study. Explain why the change addresses reporting, while recognizing that no collection mode guarantees 0 error0\text{ error}.

Explain a possible effect in context

A label such as “undercoverage” is only the beginning. A complete explanation connects the collection detail to the relevant difference and then to the estimate.

Build the explanation one link at a time
Collection detailOnly nearby students are on the frame

Farther-away students cannot be chosen.

Relevant differenceFarther-away students have longer journeys

This relationship is supplied in the example.

Possible effectSchool-wide mean may be underestimated

The excluded higher values would raise the target mean.

This is the undercoverage example’s reasoning, not a rule that all undercoverage lowers estimates.

Writing frame: “This procedure has a potential source of [type] bias because [specific selection, response or measurement detail]. [Excluded, overrepresented or misreporting group] may have [relevant difference]. As a result, [named statistic] may [overestimate / underestimate] [named population parameter].”

Adapt the frame to the evidence. Say “may” or “could” when assessing potential bias. If the description supplies a clear tendency, explain it; if it does not, say the direction cannot be determined from the information given.

Make the target and direction explicit.
Incomplete answerStronger answer
“It is biased because some people are missing.”“The parking-permit frame excludes students without permits. If their journeys are longer, the permit-only mean may underestimate the mean Tuesday travel time of all enrolled students.”
“Nonresponse makes the answer too low.”“In the library example, frequent borrowers are more likely to answer, so the responding members’ mean may overestimate the mean number of September books borrowed by all registered members.”
“The open poll is wrong.”“The 70%70\% is the dissatisfied proportion among voluntary poll respondents. Their self-selection means it should not automatically be used as the dissatisfied proportion among all registered members.”

When the direction is unknown

A researcher’s list excludes members who joined this month, but the prompt says nothing about their opinions. You can identify undercoverage: those members cannot be chosen. You cannot conclude whether the satisfaction proportion is too high or too low without information relating membership age to satisfaction.

When several sources are possible

An online survey can have limited coverage, self-selection and a leading question at the same time. The delivery mode alone does not identify the method: an online questionnaire sent to an SRS can be a valid collection option, while a public link creates a different selection process. Explain the action described, rather than treating “online” as a diagnosis.

Common mistakes and how to fix them

Use the study’s mechanism as evidence.
MistakeWhy it failsBetter approach
“A large sample cannot be biased.”Size does not repair missing coverage, self-selection or distorted responses.Inspect the collection process before relying on the count.
Calling every missing person nonresponse.A person absent from the frame was not selected and then lost.Distinguish availability for selection from failure to answer after selection.
Calling a refusal response bias.A refusal supplies no requested answer; response bias concerns distortion in supplied values.Check whether the data are absent or systematically inaccurate.
“Everyone agreed to answer, so it is voluntary response sampling.”Consent does not describe how people entered the sample.Find whether selection preceded consent or people opted into an open invitation.
“Every kind of bias lowers the result.”Different mechanisms can push estimates upward or downward.Connect the group difference or reporting tendency to the specific mean or proportion.
“Only 40%40\% replied, so upward bias is proven.”The response rate does not tell you the nonrespondents’ values.Explain how response behavior is related to the study variable, if that information is available.
“The sample mean differs from the population mean, so the method is biased.”One difference can result from chance variation.Look for a systematic feature of the procedure, not merely one inaccurate estimate.
“Random selection or a census eliminates all error.”Missing responses and distorted measurements can remain even after good selection, or when everyone is targeted.Inspect coverage, actual response and measurement as separate steps.

Eight practice questions with hints and solutions

These original fictional scenarios ask you to diagnose the process, explain a possible effect or suggest a relevant improvement. Use the given evidence rather than guessing unobserved values.

1. Identify the missing frame members

A college wants mean commute time for all enrolled students. It samples only from campus parking-pass holders. The prompt states that excluded bus riders tend to have longer commutes than pass holders. Identify the concern and explain a likely direction.

Hint

Were bus riders chosen and then lost, or were they outside the selection pool?

Solution and explanation

Undercoverage. Students without parking passes cannot be selected from the frame. Because excluded bus riders tend to have longer commutes, the pass-holder sample mean may underestimate mean commute time for all enrolled students. Use a complete enrollment frame and an appropriate random selection process.

2. An open poll about dissatisfaction

A café asks all website visitors, “Had a bad experience? Complete our optional survey.” It uses the responding visitors’ dissatisfied proportion as an estimate for all customers. What selection problem is present, and why might the estimate be high?

Hint

Who chooses which visitors become respondents? Which visitors does the invitation attract?

Solution and explanation

The sample is voluntary response: visitors choose themselves into the optional poll. Those with bad experiences may be especially motivated by the invitation, so the dissatisfied proportion among respondents could overestimate the dissatisfied proportion among all customers. The invitation’s wording is an additional concern, and customers who never visit the website may also be missed.

3. Selected members do not answer

A recreation club selects an SRS of 150150 members. Only 6060 respond to a question about weekly exercise hours. No information is provided about how respondents and nonrespondents differ. Identify the concern, calculate the simple response fraction, and state whether its direction is known.

Hint

Separate the response count from the content of the missing answers.

Solution and explanation

There is a potential nonresponse bias concern. The response fraction is 60150=40%\frac{60}{150}=40\%; 9090 selected members did not answer. Bias could arise if their exercise hours differ systematically from respondents’ hours. The direction cannot be determined from the count alone.

4. Classify the recorded-answer problem

Every student in an SRS answers a survey. Because answers are read aloud to classmates, the prompt states that students systematically report fewer hours of weekend gaming than they actually played. Identify the concern and the direction for estimated mean gaming time.

Hint

The answers are present. What happened to their accuracy?

Solution and explanation

Response bias. The collection setting encourages systematic underreporting of gaming hours. The mean reported gaming time may underestimate the true population mean. A suitable private or confidential response setting could reduce this pressure; random selection alone does not fix it.

5. Convenience versus voluntary response

Plan A surveys the first 4040 students a researcher sees near a cafeteria. Plan B posts an open link and includes anyone who chooses to submit it. Identify the selection method in each plan. Is either described as a random sample?

Hint

In one plan, easy access determines the group. In the other, participants opt in.

Solution and explanation

Plan A is convenience sampling; the researcher uses accessible students. Plan B is voluntary response sampling; people choose themselves into the sample. Neither describes a chance mechanism selecting students from a population frame. Both can produce potential bias, through different selection processes.

6. Rewrite a leading question

A school asks, “Don’t you agree our excellent new lunch service should continue?” Rewrite it to investigate student support without suggesting a preferred answer. Does better wording alone fix a survey restricted to students already in the lunch club?

Hint

Remove praise and pressure. Then consider which students can be selected.

Solution and explanation

One option is: “Do you support or oppose continuing the new lunch service?” Offer balanced options such as support, oppose and no opinion, with an appropriate “not familiar with the service” option. This addresses leading wording. It does not fix the incomplete lunch-club frame if the target is all enrolled students; use a selection process covering that target.

7. Assess the “more responses” fix

A public online poll about a local park attracts 6,0006{,}000 volunteers. A researcher says, “That many answers make bias impossible.” Assess the claim and describe a relevant improvement.

Hint

Does the larger count change how people entered the poll?

Solution and explanation

The claim is incorrect. A large count does not remove the self-selection mechanism or ensure that volunteers represent all members of the target population. Define the target, use a frame covering it, choose an appropriate random sample, contact selected units and follow up. Neutral questions and careful measurement address other stages. The prompt does not give enough information to assign a direction or exact size to the poll’s potential bias.

8. Bias or one sample’s chance error?

A hypothetical population has a known mean of 25 min25 \,\mathrm{min}. A carefully conducted SRS with complete response gives a mean of 27 min27 \,\mathrm{min}. Does this single result prove that the procedure is biased? Explain what you can conclude.

Hint

Bias describes a systematic tendency, not just one difference.

Solution and explanation

No. This estimate is 2 min2 \,\mathrm{min} above the parameter, but that difference can occur by chance in random sampling. It does not establish that the method consistently overestimates the mean. Inspect the selection and measurement procedure for systematic problems. A complete response count does not, by itself, verify measurement accuracy.

Quick revision checklist

BiasA systematic tendency to miss the target in one direction.
Chance variationEstimates move between random samples; one miss does not prove bias.
UndercoverageA relevant part of the population is absent or poorly included in the selection process.
Voluntary responsePeople choose themselves into an open invitation’s sample.
NonresponseSome selected units do not provide data; relevant differences may distort results.
Response biasSupplied answers or measurements are systematically inaccurate.
Convenience samplingChoose accessible units rather than using a population-based chance process.
DirectionUse the relevant group difference; admit when it is unknown.

Quick questions students often ask

Is every online survey a voluntary response sample?

No. “Online” describes a collection mode. A questionnaire sent to randomly selected people is different from a public poll anyone can choose to enter. Examine the selection process and actual response, not just the device used.

Does a low response rate prove bias?

No. It increases concern about missing data, but a rate does not describe the missing values. Nonresponse bias depends on relevant differences between respondents and nonrespondents and how those differences affect the estimate.

Can a census have response problems?

Yes. A complete census obtains relevant data from every member of the population, so there is no variability from choosing a sample. A census attempt can still miss units or receive unanswered items; even complete coverage does not remove distorted measurements. Attempting to ask everyone does not guarantee accurate data from everyone.

Does bias always mean deliberate dishonesty?

No. A researcher may use an outdated list, people may misunderstand an item, or an instrument may be miscalibrated without anyone intending to mislead. Explain the statistical mechanism rather than assigning motives.

Final understanding check

A fictional school wants mean Tuesday travel time for all 1,2001{,}200 students. It uses a current complete enrollment list and selects an SRS of 120120 students. 48 answer48\text{ answer} 1 invitation1\text{ invitation} asking, “Your journey is normally short, under 20 min20 \,\mathrm{min}, right?” The prompt states that students with longer journeys are less likely to answer and that the wording encourages people to describe their journeys as short.

  1. Identify the selected sample size and responding count.
  2. Explain two distinct sources of possible bias using the given details.
  3. State a likely direction for the reported picture of journey length.
  4. Explain why taking another 120120 students with the same collection process does not automatically solve the problem.
  5. Suggest a targeted improvement for each of the two concerns.
Reveal the full solution

1. Counts: 120120 students are selected; 4848 respond. The simple response fraction is 48120=40%\frac{48}{120}=40\%; 7272 selected students do not answer. Do not confuse the planned sample with the responding group.

2. Nonresponse: selected students with longer journeys are less likely to answer. Their missing data can make the responding group overrepresent shorter journeys.

Response bias and measurement mismatch: the wording suggests a short-time answer and asks about “normally,” rather than the specified Tuesday. It also collects a threshold judgment instead of the numeric duration needed to calculate a mean. The resulting answers do not directly supply the study’s intended variable.

3. Direction: given the stated response pattern and pressure toward “short,” the reported picture may make journeys look shorter than they are in the full population. However, this yes/no item cannot directly produce a mean travel time; even a reported mean would need actual duration measurements.

4. Size: choosing more students while retaining the same 11-invitation and leading-question process leaves those mechanisms in place. More answers alone do not restore missing long journeys or collect the intended durations.

Repair the stage where each problem occurs
Keep appropriate selectionComplete frame + SRS

The described enrollment frame and random selection are suitable starting points.

Improve participationFollow up with selected students

Use suitable times and response options, especially for the originally selected students who did not answer.

Improve measurementAsk for Tuesday minutes neutrally

Define leaving home, arriving at school and the actual date. Avoid suggesting a preferred duration.

5. Improvements: use follow-up to address the nonresponse concern, and replace the leading threshold item with a neutral numeric travel-time question tied to the actual Tuesday. A brief pilot can check understanding. These changes reduce identified risks; they do not guarantee that every answer will be accurate.

Coverage check: undercoverage is not indicated by the stated frame, because the current enrollment list is complete. Do not add that diagnosis merely because only 4848 students respond. Their missing classmates were selected and then did not answer.

Ready to move on? You should be able to locate a problem in the selection or collection process, explain why it matters for the variable, and distinguish a supported direction from an unsupported guess.

Continue learning

Previous: Topic 1.11 — Random Sampling · Review the sources of bias · Review contextual explanations · Back to the lesson overview