AP Statistics / Unit 1: Exploring One-Variable Data and Collecting Data / Topic 1.11
NUM8ERS study notes · Topic 1.11

Random Sampling

How can a smaller group help us learn about a whole population? Learn how to choose a sample using chance, recognize four random sampling methods, and explain which method fits a study.

2026–27 curriculum4 worked examples8 practice questionsPlans + visual comparisons

By the end of this lesson, you should be able to:

  • Distinguish sampling with replacement from sampling without replacement.
  • Explain what makes a sample a simple random sample, or SRS.
  • Recognize simple random, stratified, cluster and systematic random sampling.
  • Write a selection procedure with labels, a random mechanism and a clear stopping rule.
  • Justify a sampling method using the question, population and practical constraints.

Before you start: Know the population, sample and observational unit in a study. Review Topic 1.10 for the difference between random selection and random assignment.

First time learning this? Start with the school example and compare the four selection plans. Then work through the numbered examples.

Here to revise? Use the method checklist, then try the practice questions before opening their solutions.

The concept in 60 seconds

Random sampling uses a chance mechanism to decide who or what enters a sample. It replaces personal choice with a specified selection process. “Pick people who look typical” and “ask whoever is nearby” do not describe random sampling.

Keep the three groups clear
PopulationThe complete group the question concerns.
Sampling frameThe list or organized set from which units can be selected.
Selected sampleThe units chosen using the stated procedure.

Example: all enrolled students → complete enrollment list → selected student IDs. A missing student on the list cannot be selected from it.

Different random methods select individuals, select within every subgroup, select whole groups, or follow a fixed interval after a random start. The action that creates the sample tells you the method.

Selection and assignment do different jobs. Random selection supports learning about the population from which units were selected. Random assignment places participants into treatment groups. A random sample of students whose existing travel methods are recorded is still an observational study.

A chance procedure does not guarantee that one particular sample perfectly matches its population. Good coverage, appropriate measurement and obtaining responses still matter. This lesson focuses on the selection plan; Topic 1.12 examines sampling problems in more detail.

Cedar High wants to estimate travel time

In this fictional study, Cedar High has 1,2001{,}200 currently enrolled students. It wants to estimate their mean one-way home-to-school travel time, in minutes, on a specified Tuesday. The school plans to select 120120 students and collect each selected student’s travel time using the same measurement rule.

The school has a complete student list with grade levels. It also has 4040 advisory groups of 3030 students each; every student belongs to exactly 1 group\text{exactly }1\text{ group}. For the cluster option below, suppose the advisory groups mix grade levels and travel arrangements reasonably well.

Four possible plans for the same school. All numbers and arrangements are invented teaching examples.
MethodWhat gets selected?A possible plan
Simple randomIndividuals from the whole list.Select 120120 distinct student IDs at random from all 1,2001{,}200.
Stratified randomIndividuals from every grade level.Take a separate SRS in each grade, totaling 120120 students.
Cluster randomWhole advisory groups.Take an SRS of 44 of the 4040 groups; collect data from all 3030 students in each selected group.
Systematic randomPositions on an ordered list.Choose a starting position at random from 1–101\text{-}10, then select every 10th10^{\mathrm{th}} student on the complete list.

These plans can all use chance, but they do not allow the same sets of 120120 students. A cluster plan includes whole selected groups; a stratified plan includes a specified number from every grade.

Think first: if the school takes 3030 volunteers from each of 44 grades, does including every grade make the sample stratified random?

Check your prediction

No. The groups are defined, but students volunteer rather than being randomly selected within each grade. A stratified random plan requires a random selection process in every stratum. Matching subgroup counts alone does not supply that process.

Simple random sampling and replacement

What makes a sample an SRS?

For a simple random sample of nn distinct units from a finite population, every possible set of nn units has the same chance of being selected. In the examples on this page, SRS means selection without replacement unless a different rule is stated.

Give every unit 1 unique label1\text{ unique label}, then use a procedure that treats labels equally: for example, a uniform random-number generator or well-mixed, identical numbered slips drawn without returning them. An alphabetical list is fine for assigning labels; it is the selection mechanism that must be random.

Equal individual chances do not by themselves establish an SRS

Population: 44 students, A, B, C and D. Select 22 students.

SRS: all 66 pairs possible
AB\mathtt{AB}AC\mathtt{AC}AD\mathtt{AD}BC\mathtt{BC}BD\mathtt{BD}CD\mathtt{CD}

Each pair has chance 16\frac16.

Fair coin: choose AB\mathtt{AB} or CD\mathtt{CD}
AB\mathtt{AB}CD\mathtt{CD}

Each student has chance 12\frac12 of inclusion, but AC\mathtt{AC}, AD\mathtt{AD}, BC\mathtt{BC} and BD\mathtt{BD} are impossible.

The coin procedure is random but is not an SRS of 22 students. Its possible samples are restricted.

Without replacement: each unit can appear once

Once a unit is selected, it is no longer eligible for another selection in that sample. Drawing 44 slips and keeping them out gives 44 distinct units. With a random-number generator that can repeat values, ignore repeated labels and continue until you have the required number of distinct valid labels.

With replacement: a unit can be selected again

After each draw, return the selected unit’s label to the selection pool. A repeated label is allowed and counts as another draw. For example, the 4-draw4\text{-draw} sequence 03,  11,  03,  08\mathtt{03},\;\mathtt{11},\;\mathtt{03},\;\mathtt{08} contains 44 draws but only 33 distinct units.

Without replacement

Pool gets smaller after a selection. Do not count the same unit twice.

For 44 distinct students, 03,  11,  03,  08\mathtt{03},\;\mathtt{11},\;\mathtt{03},\;\mathtt{08} is incomplete: skip the repeated 03\mathtt{03} and keep selecting.

With replacement

Pool is restored after each draw. Repeated labels are permitted.

The repeated 03\mathtt{03} remains valid in this 4-draw4\text{-draw} sequence. Do not silently change the rule by discarding it.

Replacement concerns the selection process. It does not mean “replace a selected person who does not answer with any available volunteer.” That changes how the responding group is formed.

How to describe an SRS precisely

  1. Define the frame: use the complete, current list of the population’s units.
  2. Label once: assign a different label to every unit.
  3. Name the random mechanism: select labels uniformly using a random-number generator, or draw well-mixed identical slips.
  4. State the replacement rule: for distinct units, select without replacement or ignore repeats.
  5. Stop at the required size: collect data from the units corresponding to the accepted labels.

For a random-digit table, use fixed-width labels and read non-overlapping groups of that width. State a starting point and direction. Ignore numbers outside the label range and, for sampling without replacement, labels already accepted. Do not scan the table for numbers you prefer.

Stratified, cluster and systematic random sampling

Stratified random: take an SRS from every stratum

Divide the entire population into non-overlapping groups called strata, using a characteristic available before selection. Every unit must belong to 1 stratum1\text{ stratum}. Then select an SRS within each stratum and combine those selected units.

Useful strata often contain units similar with respect to the response, while different strata have different response patterns. For travel time, residential areas or grade levels might be helpful if there is a reason to expect travel times to differ across them. Grouping by a completely unrelated characteristic may bring little improvement.

Stratification ensures that each stratum contributes selected units when a positive sample size is allocated to it. It can improve the precision of an estimate when the grouping is informative. It does not guarantee every selected unit responds.

Cluster random: select groups, then include everyone in them

Divide the population into non-overlapping groups called clusters. Select an SRS of the clusters and collect data from all units in the selected clusters. The one-stage cluster design taught here takes no units from unselected clusters.

Ideally, each cluster contains a mix resembling the population, and different clusters are reasonably similar. In practice, existing groups such as schools, neighborhoods or classes can be convenient to reach. Their members may be very similar to one another, so cluster sampling can give less precise estimates than a well-spread individual sample of the same size.

Stratified versus cluster: watch which groups contribute

Tiny schematic populations: each numbered tile is 1 unit1\text{ unit}; a check mark means selected. These illustrate selection, not collected outcomes.

Stratified: 22 from each of 33 strata
Stratum A · 88 units
Stratum B · 88 units
Stratum C · 88 units

Selected: A02\mathtt{A02}, A06\mathtt{A06}; B01\mathtt{B01}, B05\mathtt{B05}; C03\mathtt{C03}, C08\mathtt{C08}. Every stratum contributes 22 units.

Cluster: all 66 in 11 of 44 clusters
Cluster 11 · 66 units
Cluster 22 · 66 units · selected
Cluster 33 · 66 units
Cluster 44 · 66 units

Selected: units 2.1–2.6\mathtt{2.1}\text{–}\mathtt{2.6}. Only Cluster 22 contributes; every unit in it is included.

Stratified: some units from every group. Cluster: all units from selected groups. The group names alone do not identify the method.

If you randomly choose groups and then randomly choose only some units inside them, that is a multistage design. It differs from the one-stage cluster plan above. This distinction is useful for recognizing a description; detailed multistage analysis is beyond this lesson.

Systematic random: random start, then a fixed interval

Choose a starting position randomly from the first kk positions, then select every kthk^{\mathrm{th}} unit on the ordered frame. When NN is an exact multiple of nn, the interval can be k=Nnk=\frac{N}{n}. Use an example with a whole-number interval unless a more detailed rule is provided.

A random start determines the rest of a systematic sample

Illustration: 4040 positions, sample size 44, interval 1010. Suppose the uniform random start from 1–101\text{-}10 is 77.

Selected positions: 77, 1717, 2727, 3737. The check marks repeat every 1010 positions. The starting value is an illustrative possible outcome.

Starting automatically at the first unit is systematic selection, but it does not include the random-start feature needed for a systematic random sample. On a fixed list, systematic sampling usually permits far fewer possible samples than an SRS of the same size.

Check the order of the list. If a repeating list pattern matches the selection interval, a systematic sample may repeatedly hit just one position in that pattern. For example, selecting every 10th10^{\mathrm{th}} production record can miss important variation if each block of 1010 is arranged by the same machine order. A random start alone does not remove this concern.

A systematic design with a uniformly random start can give all units equal inclusion chances in the exact-multiple example. That still does not make it an SRS: equal chances for units and equal chances for all possible samples are different requirements.

Choose a method and justify the choice

No sampling method is best for every study. Consider what information is available before sampling, which subgroups need representation, how far apart units are, and how data will be collected. A useful justification connects those features to the method.

Reasons to choose a method, with the tradeoff that should stay in view.
MethodA reason it may fitA limitation to consider
SRSA complete individual list is available, and no known subgroup information needs to guide selection.Small subgroups may contribute few selected units by chance. In-person visits may be widely scattered.
Stratified randomImportant subgroups must contribute, or the strata are expected to differ in the response.You need reliable stratum membership before sampling. Different selection fractions can require weights.
Cluster randomReaching entire nearby groups costs less than reaching individuals scattered across the population.Similar units within clusters may reduce precision. Unequal cluster sizes can produce a variable number of sampled units.
Systematic randomAn ordered list or stream is available, and a regular selection interval is easy to carry out.The frame order must be examined for repeating patterns related to the response.

Allocate a stratified sample deliberately

Cedar High’s 44 grade sizes are 360360, 300300, 300300 and 240240. For a sample of 120120, a proportional allocation selects 10%10\% from each grade: 3636, 3030, 3030 and 2424. Each selection is an SRS within that grade.

Population size → selected sample size in each grade
Grade 9360→36360\to36
Grade 10300→30300\to30
Grade 11300→30300\to30
Grade 12240→24240\to24

36+30+30+24=12036+30+30+24=120. Each selected fraction is 0.100.10. The counts are a planned allocation, not observed travel-time data.

Proportional allocation is one option, not part of the definition of stratified random sampling. A researcher may deliberately select more units from a small stratum to study it separately. If selection fractions differ, do not automatically pool every response equally to estimate the whole-population mean; the analysis must reflect the population shares.

Example of why weights can matter: if 2 strata2\text{ strata} contain 900900 and 300300 students, sampling 3030 from each makes 12 of the sample\frac12\text{ of the sample} come from the smaller stratum even though it contains 14 of the population\frac14\text{ of the population}. For the population mean, the 2 stratum means2\text{ stratum means} would be combined using population shares 0.750.75 and 0.250.25, rather than 0.500.50 and 0.500.50. Detailed survey estimation is beyond this lesson.

For the school question, stratifying by grade is justified if grade-specific travel patterns or guaranteed grade representation matter. If students’ travel times do not meaningfully differ by grade, grade stratification may provide little precision gain. Explain the reason instead of simply saying “stratified is more accurate.”

Four worked examples

All scenarios and possible random outputs below are invented for teaching. The selected labels show how to apply a rule; they are not findings from real surveys.

Example 1: Select an SRS using random digits

Task: a club has 4848 registered members. Select 44 distinct members for a survey using a random-digit table.

  1. Use the complete current membership list. Assign each member exactly 1 label\text{exactly }1\text{ label} from 01 to 48\mathtt{01}\text{ to }\mathtt{48}.
  2. Specify a starting location and reading direction before reading. Read successive non-overlapping 2-digit2\text{-digit} groups from a table of random digits.
  3. Accept labels 01–48\mathtt{01}\text{–}\mathtt{48}. Ignore 00\mathtt{00} and 49–99\mathtt{49}\text{–}\mathtt{99}. Because selection is without replacement, ignore a label already accepted.
  4. Stop after 44 distinct valid labels, then survey their corresponding members.

Suppose the successive 2-digit2\text{-digit} groups at the specified location are:

07,  58,  07,  00,  31,  18,  05\mathtt{07},\;\mathtt{58},\;\mathtt{07},\;\mathtt{00},\;\mathtt{31},\;\mathtt{18},\;\mathtt{05}

Apply the rule in order; do not choose which numbers to keep.
Group readDecisionReason
07\mathtt{07}AcceptValid, not previously selected.
58\mathtt{58}SkipOutside 01–48\mathtt{01}\text{–}\mathtt{48}.
07\mathtt{07}SkipAlready selected; no replacement.
00\mathtt{00}SkipNo member has label 00\mathtt{00}.
31,  18,  05\mathtt{31},\;\mathtt{18},\;\mathtt{05}Accept each, in orderAll are valid and distinct. Stop after 05\mathtt{05}.

Selected sample: members 07,  31,  18,  05\mathtt{07},\;\mathtt{31},\;\mathtt{18},\;\mathtt{05}. Skipping invalid groups and duplicates preserves the stated no-replacement rule when the digit groups come from a suitable random source.

A uniform random-number generator selecting 44 distinct integers from 1–481\text{-}48 is another valid procedure. Be explicit that repeats are prevented or ignored; naming a calculator command without explaining its settings is less clear.

Example 2: Build a proportional stratified sample

Task: Cedar High wants all 44 grades represented in its 120120-student travel-time sample. The grade sizes are 360360, 300300, 300300 and 240240.

  1. Define strata: the 44 grade levels. Every enrolled student belongs to exactly 1 grade\text{exactly }1\text{ grade}.
  2. Allocate: select 3636, 3030, 3030 and 2424 students respectively. For example, Grade 9 receives 3601,200×120=36\frac{360}{1{,}200}\times120=36 selections.
  3. Randomize within each grade: give students unique labels on that grade’s complete list and use a uniform generator to choose the required number of distinct labels.
  4. Combine: survey the selected students from all 44 grades, using the same travel-time definition and specified Tuesday.

Method: stratified random sampling, because an SRS is taken from every grade. Each grade’s selection fraction is 10%10\%.

Justification: this plan guarantees selected representation from each grade and can improve precision if travel patterns differ by grade. The 44 grade samples combined are not an SRS of 120120 from the entire school: sets with the wrong grade counts are not possible under this allocation.

Example 3: Reach fewer groups with cluster sampling

Task: the school can collect travel-time responses efficiently during advisory meetings. Its 4040 advisory groups each contain 3030 students, and the groups are intended to contain similar mixes of grades and travel arrangements.

  1. Use a complete list of all 4040 non-overlapping advisory groups, covering all 1,2001{,}200 students.
  2. Label groups 01–40\mathtt{01}\text{–}\mathtt{40}. Use a uniform random-number generator to select 44 distinct group labels.
  3. Suppose the selected labels are 04,  11,  27,  38\mathtt{04},\;\mathtt{11},\;\mathtt{27},\;\mathtt{38}. Collect the relevant response from every student in those groups.
  4. The planned sample contains 4×30=120 students4\times30=120\text{ students}, located in 44 selected groups.

Method: one-stage cluster random sampling. The groups are selected randomly; their individual students are all included.

Justification: collecting in 4 meetings4\text{ meetings} can be simpler than reaching students spread across all 4040 meetings. The tradeoff is that students within the same advisory group may share travel patterns, so 120120 clustered responses need not give the same precision as 120120 individually sampled responses.

Random selection of groups supports population inference with appropriate collection and analysis; it does not guarantee these 44 groups perfectly mirror the whole school. If groups had unequal sizes, 44 selected groups would not necessarily give exactly 120120 students.

Example 4: Specify a systematic random sample

Task: select 120120 positions from Cedar High’s complete 1,2001{,}200-student list using a regular interval.

  1. Check the frame order: identify any repeating arrangement related to travel time before using this plan.
  2. Calculate the interval: k=1,200120=10k=\frac{1{,}200}{120}=10.
  3. Randomize the start: use a uniform random generator to choose 1 integer1\text{ integer} from 1–101\text{-}10. Suppose it is 77.
  4. Follow the interval: select positions 77, 1717, 2727, 3737, …, 1,1971{,}197.
  5. Check the count: the last position is 7+119×10=1,1977+119\times10=1{,}197, so the plan selects 120120 positions.

Method: systematic random sampling, because a random start is followed by a fixed interval.

Justification: the plan is easy to carry out on an ordered list and spreads selections along it. Its suitability depends on the list’s order. It is not an SRS of 120120, because most possible sets of 120120 students cannot be formed by taking every 10th10^{\mathrm{th}} position.

Do not start by choosing any number from 1–1,2001\text{-}1{,}200 and then simply stop at the end: that procedure may fail to select 120120 students and is not the stated plan.

Write a complete sampling plan in context

A reader should be able to follow your procedure without making a new selection decision. “Choose students randomly” identifies an intention, but leaves the actual method unclear.

Population and frameName the group and the complete list or group structure used for selection.
Unit and labelsState whether labels identify people, objects, groups or list positions.
Random mechanismDescribe a uniform random generator or another suitable chance procedure.
Grouping or intervalSpecify every stratum’s sample, selected clusters, or the systematic start and interval.
Replacement and stoppingExplain duplicates, invalid labels and when selection ends.
Collection and reasoningSay whose responses are obtained, what is measured, and why the method fits.

Procedure frame: “Use [complete frame] for [defined population]. Label [units] from [range]. Use [random mechanism] to select [number] [distinct labels / groups / starting position]. [State the within-group or interval rule.] Collect [information] from [selected units].”

Justification frame: “This method fits because [a feature of this population or collection setting] makes [a specific design feature] useful. A limitation is [relevant tradeoff].”

Name evidence, not just a label

A complete identification sounds like: “This is stratified random sampling because the school divides students by grade and takes an SRS from each grade.” For cluster sampling, state that all students in selected groups are included. For systematic sampling, identify both the random start and the interval.

Keep the responding sample in view

The selected sample and the group that actually supplies data may differ. If a randomly selected student does not respond, inviting a convenient replacement is not automatically the same design. Plan an appropriate contact and follow-up process; distinguish missing data from a response of 0\text{response of }0.

Use student IDs in the working selection record and collect only the information the question needs. The examples here require no private student information. Sampling labels identify selection units; they are not measurements of travel time.

Finally, keep the scope tied to the population covered by the frame. A complete Cedar High enrollment list supports selection for that school, not all teenagers. Sampling does not turn an observational association into a causal treatment effect.

Common mistakes and how to fix them

Use the described selection actions as your evidence.
MistakeWhy it failsBetter approach
“Everyone has an equal chance, so it must be an SRS.”A design can restrict which groups of units are possible.Check whether every possible sample of the required size has the same chance.
Calling a sample random because nobody deliberately chose the names.An arbitrary or convenient choice need not use chance.Identify the actual random mechanism.
Calling “some from every group” cluster sampling.One-stage cluster sampling includes all units in selected groups.For stratified random sampling, take an SRS from each stratum.
Calling every group-based plan stratified random.The method depends on how groups and units are selected.Distinguish samples from every group, whole selected groups, and multistage selection.
Keeping a repeated label in a sample of distinct units.It counts the same unit twice under a no-replacement rule.Skip the repeat and continue until the required distinct count is reached.
Starting every kthk^{\mathrm{th}} selection at position 11 without randomizing.The random-start feature is missing.Choose uniformly among positions 1–k1\text{–}k before following the interval.
Claiming stratified samples must use equal counts per group.Allocation can be proportional or chosen for another justified goal.State each stratum’s size and selection rule; consider weighting when fractions differ.
“Random sampling guarantees a representative sample and proves a cause.”Chance variation remains; selection does not assign treatments.Explain population coverage separately from experimental assignment.

Eight practice questions with hints and solutions

These are original fictional scenarios. Give the method or calculation requested and explain it using the actual selection process.

1. Identify an SRS

A library has a complete list of 400400 registered members. Each receives a unique label 001–400\mathtt{001}\text{–}\mathtt{400}. A uniform random generator selects 2020 distinct labels, and those members are surveyed. Identify the sampling method and replacement rule.

Hint

Are labels chosen from the whole frame? Can a label be counted twice?

Solution and explanation

SRS without replacement. The generator chooses 2020 distinct labels uniformly from the complete membership frame. Every set of 2020 members has the same selection chance, and no member is selected twice.

2. Equal individual chances, restricted samples

A fair coin selects either students A and B or students C and D. Each student has a 12\frac12 chance of inclusion. Is this an SRS of 22 students from the 4\text{from the }4?

Hint

Could A and C ever form the selected pair?

Solution and explanation

No. Only AB\mathtt{AB} and CD\mathtt{CD} can be selected. AC\mathtt{AC}, AD\mathtt{AD}, BC\mathtt{BC} and BD\mathtt{BD} have selection chance 0\text{selection chance }0, while AB\mathtt{AB} and CD\mathtt{CD} each have chance 12\frac12. Equal individual inclusion chances do not give equal chances to all 66 possible pairs.

3. Apply the replacement rule

From labels 01–30\mathtt{01}\text{–}\mathtt{30}, 4 successive4\text{ successive} uniform random draws give 06,  21,  06,  18\mathtt{06},\;\mathtt{21},\;\mathtt{06},\;\mathtt{18}. Compare the interpretation with replacement and without replacement when 44 distinct students are required.

Hint

Count the draws separately from the distinct labels.

Solution and explanation

With replacement: all 44 draws are valid, including the repeated 06\mathtt{06}; there are 33 distinct students. Without replacement: accept 06,  21,  18\mathtt{06},\;\mathtt{21},\;\mathtt{18}, skip the repeated 06\mathtt{06}, and continue drawing until a 4th distinct4^{\mathrm{th}}\text{ distinct} valid label is accepted.

4. Allocate within strata

A school’s 22 grade groups contain 900900 and 300300 students. It wants a proportional stratified random sample of 8080. How many students should it select from each grade? Would selecting 4040 at random from each still be stratified random sampling?

Hint

Multiply each population share by 8080. Separate the method from its allocation.

Solution and explanation

The population total is 1,2001{,}200. The allocations are 9001,200×80=60\frac{900}{1{,}200}\times80=60 and 3001,200×80=20\frac{300}{1{,}200}\times80=20. An SRS of 4040 from each grade would still be stratified random sampling, but it would not be proportional. It overrepresents the smaller grade in the unweighted sample, so whole-school estimation must account for the unequal population shares.

5. Select whole clusters

A camp has 2020 cabins, each with 66 campers. Every camper belongs to 1 cabin1\text{ cabin}. The director selects an SRS of 33 cabins and surveys every camper in them. Identify the sampling method and planned number of sampled campers. Is it a census of the camp?

Hint

What is randomly selected, and who inside the selected groups is included?

Solution and explanation

One-stage cluster random sampling. Cabins are randomly selected clusters, and all campers in them are included. The planned sample is 3×6=18 campers3\times6=18\text{ campers}. It is not a census of the entire camp’s 120120 campers, even though every camper in the 3 selected cabins3\text{ selected cabins} is surveyed.

6. Find the interval and endpoint

A complete list contains 240240 items. A systematic random sample will contain 2424 items. Determine the interval and random-start range. If the start is 44, give the first 3 positions\text{first }3\text{ positions} and the last position.

Hint

Use k=Nnk=\frac{N}{n}, then add kk between selections. There are 2323 steps after the first selection.

Solution and explanation

k=24024=10k=\frac{240}{24}=10. Choose the start uniformly from 1–101\text{-}10. Starting at 44 gives 44, 1414 and 2424 initially. The last position is 4+23×10=2344+23\times10=234. This is a systematic random sample, not an SRS of 2424, because the interval restricts possible samples.

7. Justify a method

A library system has 44 branches. Each member is registered at exactly 1 branch\text{exactly }1\text{ branch}, complete member lists are available, and borrowing patterns are expected to differ across branches. The researcher wants all 44 branches to contribute to a member survey. Suggest and justify a random sampling method.

Hint

Which method samples individuals within every defined subgroup?

Solution and explanation

Use stratified random sampling by registered branch. Assign a positive sample size to each branch and take an SRS without replacement from each branch’s complete member list. This gives selected representation from all 44 branches and may improve precision because borrowing patterns differ across branches. Specify allocation and account for population shares if using unequal selection fractions.

8. Read random digits with a stopping rule

30 members30\text{ members} are labeled 01–30\mathtt{01}\text{–}\mathtt{30}. Select an SRS of 44 distinct members by reading these successive 2-digit2\text{-digit} groups: 31,  05,  00,  05,  14,  22,  30,  09\mathtt{31},\;\mathtt{05},\;\mathtt{00},\;\mathtt{05},\;\mathtt{14},\;\mathtt{22},\;\mathtt{30},\;\mathtt{09}. Which labels are accepted, and where do you stop?

Hint

Skip invalid labels and repeats. Stop after 44 distinct valid labels, not 44 groups read.

Solution and explanation

Skip 31\mathtt{31}, accept 05\mathtt{05}, skip 00\mathtt{00}, skip the repeated 05\mathtt{05}, then accept 1414, 2222 and 3030. Stop at 30\mathtt{30}: the 4 selected labels4\text{ selected labels} are 05,  14,  22,  30\mathtt{05},\;\mathtt{14},\;\mathtt{22},\;\mathtt{30}. Do not include 09\mathtt{09} after the target is reached.

Quick revision checklist

SRSEvery possible sample of the stated size is equally likely.
Without replacementA unit can be selected only once; skip repeated labels.
With replacementA unit can appear on multiple draws; repeats are allowed.
Stratified randomDivide by a characteristic and take an SRS from every stratum.
Cluster randomTake an SRS of groups and include all units in selected groups.
Systematic randomRandom start, then every kthk^{\mathrm{th}} unit. Examine list order.
Complete procedureFrame, labels, mechanism, rules and stopping point.
Useful justificationConnect subgroup information, cost or list structure to the study question.

Quick questions students often ask

Are strata always the same size?

No. Strata can have different population sizes, and their selected sample sizes can differ too. The defining feature is random selection within every stratum, not equal group sizes.

Is a class always a cluster?

A class can be used as a cluster when entire randomly selected classes are included. Taking an SRS of students from every class is stratified random sampling by class. Randomly selecting some classes and then sampling some students in them is multistage sampling. Read the actions, not just the group name.

Can systematic sampling be random without being an SRS?

Yes. Randomizing the start provides chance selection, but the fixed interval on an ordered list usually restricts the sets that can be selected. The 2-student coin example2\text{-student coin example} makes the same distinction between individual chances and possible samples.

Does random sampling fix an incomplete population list?

No. A random procedure applied to an incomplete frame cannot select units missing from that frame. It also does not fix inaccurate answers or missing responses. Those issues are developed in Topic 1.12.

Final understanding check

A fictional school has 240240 enrolled students living in 33 non-overlapping travel zones: Zone A has 120120, Zone B has 8080 and Zone C has 4040. Complete zone-specific student lists are available. Travel time is expected to differ across zones. The school wants to select 2424 students to estimate mean travel time on a specified Tuesday.

  1. Propose a sampling method and justify why it fits.
  2. Find the sample size for each zone under proportional allocation.
  3. Write a random selection procedure that produces distinct students.
  4. Explain why the combined sample is not an SRS of 2424 from all 240240 students.
  5. Assess a different plan: randomly select 1 zone1\text{ zone} and survey every student living in it. Identify its method and a possible practical drawback.
Reveal the full solution

1. Method and reason: stratified random sampling by travel zone. Sampling from every zone gives selected representation for each area and can improve precision because travel times are expected to differ across zones.

2. Allocation: 120240×24=12\frac{120}{240}\times24=12 from A; 80240×24=8\frac{80}{240}\times24=8 from B; 40240×24=4\frac{40}{240}\times24=4 from C. The total is 2424, and every zone’s selection fraction is 10%10\%.

3. Procedure: on each complete zone list, label students uniquely: A001–A120\mathtt{A001}\text{–}\mathtt{A120}, B001–B080\mathtt{B001}\text{–}\mathtt{B080} and C001–C040\mathtt{C001}\text{–}\mathtt{C040}. Independently within each zone, use a uniform random generator to select 1212, 88 or 44 distinct valid numbered labels respectively. Prevent or ignore repeats within that zone, continuing until its target is reached. Combine the selected students and collect the same defined Tuesday travel-time response from each.

240240 enrolled students3 complete, non-overlapping zone lists3\text{ complete, non-overlapping zone lists}
Zone A: 120120SRS of 1212
Zone B: 8080SRS of 88
Zone C: 4040SRS of 44
Combine 2424 selected studentsCollect the defined travel-time response

4. SRS distinction: the allocation always contains 1212 from A, 88 from B and 44 from C. Other sets of 2424, such as a set entirely from Zone A, have selection chance 0\text{selection chance }0. Not every possible sample of 2424 is equally likely.

5. Alternative: selecting 1 zone1\text{ zone} randomly and including every student in it is one-stage cluster random sampling using zones as clusters. Its student count is 120120, 8080 or 4040, so it does not give the target 2424. Because travel times are expected to differ between zones, collecting from only 1 zone1\text{ zone} could give a sample poorly reflecting the full range of travel patterns. The random cluster choice does not guarantee an individual sample is representative.

Ready to move on? You should be able to identify who or what is randomized, explain replacement, distinguish strata from clusters, and write a selection plan someone else could carry out.

Continue learning

Previous: Topic 1.10 — Investigative questions and data collection · Review the sampling methods · Review SRS and replacement · Back to the lesson overview