Random Sampling
How can a smaller group help us learn about a whole population? Learn how to choose a sample using chance, recognize four random sampling methods, and explain which method fits a study.
By the end of this lesson, you should be able to:
- Distinguish sampling with replacement from sampling without replacement.
- Explain what makes a sample a simple random sample, or SRS.
- Recognize simple random, stratified, cluster and systematic random sampling.
- Write a selection procedure with labels, a random mechanism and a clear stopping rule.
- Justify a sampling method using the question, population and practical constraints.
Before you start: Know the population, sample and observational unit in a study. Review Topic 1.10 for the difference between random selection and random assignment.
First time learning this? Start with the school example and compare the four selection plans. Then work through the numbered examples.
Here to revise? Use the method checklist, then try the practice questions before opening their solutions.
The concept in 60 seconds
Random sampling uses a chance mechanism to decide who or what enters a sample. It replaces personal choice with a specified selection process. “Pick people who look typical” and “ask whoever is nearby” do not describe random sampling.
Example: all enrolled students → complete enrollment list → selected student IDs. A missing student on the list cannot be selected from it.
Different random methods select individuals, select within every subgroup, select whole groups, or follow a fixed interval after a random start. The action that creates the sample tells you the method.
Selection and assignment do different jobs. Random selection supports learning about the population from which units were selected. Random assignment places participants into treatment groups. A random sample of students whose existing travel methods are recorded is still an observational study.
A chance procedure does not guarantee that one particular sample perfectly matches its population. Good coverage, appropriate measurement and obtaining responses still matter. This lesson focuses on the selection plan; Topic 1.12 examines sampling problems in more detail.
Cedar High wants to estimate travel time
In this fictional study, Cedar High has currently enrolled students. It wants to estimate their mean one-way home-to-school travel time, in minutes, on a specified Tuesday. The school plans to select students and collect each selected student’s travel time using the same measurement rule.
The school has a complete student list with grade levels. It also has advisory groups of students each; every student belongs to . For the cluster option below, suppose the advisory groups mix grade levels and travel arrangements reasonably well.
| Method | What gets selected? | A possible plan |
|---|---|---|
| Simple random | Individuals from the whole list. | Select distinct student IDs at random from all . |
| Stratified random | Individuals from every grade level. | Take a separate SRS in each grade, totaling students. |
| Cluster random | Whole advisory groups. | Take an SRS of of the groups; collect data from all students in each selected group. |
| Systematic random | Positions on an ordered list. | Choose a starting position at random from , then select every student on the complete list. |
These plans can all use chance, but they do not allow the same sets of students. A cluster plan includes whole selected groups; a stratified plan includes a specified number from every grade.
Think first: if the school takes volunteers from each of grades, does including every grade make the sample stratified random?
Check your prediction
No. The groups are defined, but students volunteer rather than being randomly selected within each grade. A stratified random plan requires a random selection process in every stratum. Matching subgroup counts alone does not supply that process.
Simple random sampling and replacement
What makes a sample an SRS?
For a simple random sample of distinct units from a finite population, every possible set of units has the same chance of being selected. In the examples on this page, SRS means selection without replacement unless a different rule is stated.
Give every unit , then use a procedure that treats labels equally: for example, a uniform random-number generator or well-mixed, identical numbered slips drawn without returning them. An alphabetical list is fine for assigning labels; it is the selection mechanism that must be random.
Population: students, A, B, C and D. Select students.
Each pair has chance .
Each student has chance of inclusion, but , , and are impossible.
The coin procedure is random but is not an SRS of students. Its possible samples are restricted.
Without replacement: each unit can appear once
Once a unit is selected, it is no longer eligible for another selection in that sample. Drawing slips and keeping them out gives distinct units. With a random-number generator that can repeat values, ignore repeated labels and continue until you have the required number of distinct valid labels.
With replacement: a unit can be selected again
After each draw, return the selected unit’s label to the selection pool. A repeated label is allowed and counts as another draw. For example, the sequence contains draws but only distinct units.
Without replacement
Pool gets smaller after a selection. Do not count the same unit twice.
For distinct students, is incomplete: skip the repeated and keep selecting.
With replacement
Pool is restored after each draw. Repeated labels are permitted.
The repeated remains valid in this sequence. Do not silently change the rule by discarding it.
Replacement concerns the selection process. It does not mean “replace a selected person who does not answer with any available volunteer.” That changes how the responding group is formed.
How to describe an SRS precisely
- Define the frame: use the complete, current list of the population’s units.
- Label once: assign a different label to every unit.
- Name the random mechanism: select labels uniformly using a random-number generator, or draw well-mixed identical slips.
- State the replacement rule: for distinct units, select without replacement or ignore repeats.
- Stop at the required size: collect data from the units corresponding to the accepted labels.
For a random-digit table, use fixed-width labels and read non-overlapping groups of that width. State a starting point and direction. Ignore numbers outside the label range and, for sampling without replacement, labels already accepted. Do not scan the table for numbers you prefer.
Stratified, cluster and systematic random sampling
Stratified random: take an SRS from every stratum
Divide the entire population into non-overlapping groups called strata, using a characteristic available before selection. Every unit must belong to . Then select an SRS within each stratum and combine those selected units.
Useful strata often contain units similar with respect to the response, while different strata have different response patterns. For travel time, residential areas or grade levels might be helpful if there is a reason to expect travel times to differ across them. Grouping by a completely unrelated characteristic may bring little improvement.
Stratification ensures that each stratum contributes selected units when a positive sample size is allocated to it. It can improve the precision of an estimate when the grouping is informative. It does not guarantee every selected unit responds.
Cluster random: select groups, then include everyone in them
Divide the population into non-overlapping groups called clusters. Select an SRS of the clusters and collect data from all units in the selected clusters. The one-stage cluster design taught here takes no units from unselected clusters.
Ideally, each cluster contains a mix resembling the population, and different clusters are reasonably similar. In practice, existing groups such as schools, neighborhoods or classes can be convenient to reach. Their members may be very similar to one another, so cluster sampling can give less precise estimates than a well-spread individual sample of the same size.
Tiny schematic populations: each numbered tile is ; a check mark means selected. These illustrate selection, not collected outcomes.
Selected: , ; , ; , . Every stratum contributes units.
Selected: units . Only Cluster contributes; every unit in it is included.
Stratified: some units from every group. Cluster: all units from selected groups. The group names alone do not identify the method.
If you randomly choose groups and then randomly choose only some units inside them, that is a multistage design. It differs from the one-stage cluster plan above. This distinction is useful for recognizing a description; detailed multistage analysis is beyond this lesson.
Systematic random: random start, then a fixed interval
Choose a starting position randomly from the first positions, then select every unit on the ordered frame. When is an exact multiple of , the interval can be . Use an example with a whole-number interval unless a more detailed rule is provided.
Illustration: positions, sample size , interval . Suppose the uniform random start from is .
Selected positions: , , , . The check marks repeat every positions. The starting value is an illustrative possible outcome.
Starting automatically at the first unit is systematic selection, but it does not include the random-start feature needed for a systematic random sample. On a fixed list, systematic sampling usually permits far fewer possible samples than an SRS of the same size.
Check the order of the list. If a repeating list pattern matches the selection interval, a systematic sample may repeatedly hit just one position in that pattern. For example, selecting every production record can miss important variation if each block of is arranged by the same machine order. A random start alone does not remove this concern.
A systematic design with a uniformly random start can give all units equal inclusion chances in the exact-multiple example. That still does not make it an SRS: equal chances for units and equal chances for all possible samples are different requirements.
Choose a method and justify the choice
No sampling method is best for every study. Consider what information is available before sampling, which subgroups need representation, how far apart units are, and how data will be collected. A useful justification connects those features to the method.
| Method | A reason it may fit | A limitation to consider |
|---|---|---|
| SRS | A complete individual list is available, and no known subgroup information needs to guide selection. | Small subgroups may contribute few selected units by chance. In-person visits may be widely scattered. |
| Stratified random | Important subgroups must contribute, or the strata are expected to differ in the response. | You need reliable stratum membership before sampling. Different selection fractions can require weights. |
| Cluster random | Reaching entire nearby groups costs less than reaching individuals scattered across the population. | Similar units within clusters may reduce precision. Unequal cluster sizes can produce a variable number of sampled units. |
| Systematic random | An ordered list or stream is available, and a regular selection interval is easy to carry out. | The frame order must be examined for repeating patterns related to the response. |
Allocate a stratified sample deliberately
Cedar High’s grade sizes are , , and . For a sample of , a proportional allocation selects from each grade: , , and . Each selection is an SRS within that grade.
. Each selected fraction is . The counts are a planned allocation, not observed travel-time data.
Proportional allocation is one option, not part of the definition of stratified random sampling. A researcher may deliberately select more units from a small stratum to study it separately. If selection fractions differ, do not automatically pool every response equally to estimate the whole-population mean; the analysis must reflect the population shares.
Example of why weights can matter: if contain and students, sampling from each makes come from the smaller stratum even though it contains . For the population mean, the would be combined using population shares and , rather than and . Detailed survey estimation is beyond this lesson.
For the school question, stratifying by grade is justified if grade-specific travel patterns or guaranteed grade representation matter. If students’ travel times do not meaningfully differ by grade, grade stratification may provide little precision gain. Explain the reason instead of simply saying “stratified is more accurate.”
Four worked examples
All scenarios and possible random outputs below are invented for teaching. The selected labels show how to apply a rule; they are not findings from real surveys.
Example 1: Select an SRS using random digits
Task: a club has registered members. Select distinct members for a survey using a random-digit table.
- Use the complete current membership list. Assign each member from .
- Specify a starting location and reading direction before reading. Read successive non-overlapping groups from a table of random digits.
- Accept labels . Ignore and . Because selection is without replacement, ignore a label already accepted.
- Stop after distinct valid labels, then survey their corresponding members.
Suppose the successive groups at the specified location are:
| Group read | Decision | Reason |
|---|---|---|
| Accept | Valid, not previously selected. | |
| Skip | Outside . | |
| Skip | Already selected; no replacement. | |
| Skip | No member has label . | |
| Accept each, in order | All are valid and distinct. Stop after . |
Selected sample: members . Skipping invalid groups and duplicates preserves the stated no-replacement rule when the digit groups come from a suitable random source.
A uniform random-number generator selecting distinct integers from is another valid procedure. Be explicit that repeats are prevented or ignored; naming a calculator command without explaining its settings is less clear.
Example 2: Build a proportional stratified sample
Task: Cedar High wants all grades represented in its -student travel-time sample. The grade sizes are , , and .
- Define strata: the grade levels. Every enrolled student belongs to .
- Allocate: select , , and students respectively. For example, Grade 9 receives selections.
- Randomize within each grade: give students unique labels on that grade’s complete list and use a uniform generator to choose the required number of distinct labels.
- Combine: survey the selected students from all grades, using the same travel-time definition and specified Tuesday.
Method: stratified random sampling, because an SRS is taken from every grade. Each grade’s selection fraction is .
Justification: this plan guarantees selected representation from each grade and can improve precision if travel patterns differ by grade. The grade samples combined are not an SRS of from the entire school: sets with the wrong grade counts are not possible under this allocation.
Example 3: Reach fewer groups with cluster sampling
Task: the school can collect travel-time responses efficiently during advisory meetings. Its advisory groups each contain students, and the groups are intended to contain similar mixes of grades and travel arrangements.
- Use a complete list of all non-overlapping advisory groups, covering all students.
- Label groups . Use a uniform random-number generator to select distinct group labels.
- Suppose the selected labels are . Collect the relevant response from every student in those groups.
- The planned sample contains , located in selected groups.
Method: one-stage cluster random sampling. The groups are selected randomly; their individual students are all included.
Justification: collecting in can be simpler than reaching students spread across all meetings. The tradeoff is that students within the same advisory group may share travel patterns, so clustered responses need not give the same precision as individually sampled responses.
Random selection of groups supports population inference with appropriate collection and analysis; it does not guarantee these groups perfectly mirror the whole school. If groups had unequal sizes, selected groups would not necessarily give exactly students.
Example 4: Specify a systematic random sample
Task: select positions from Cedar High’s complete -student list using a regular interval.
- Check the frame order: identify any repeating arrangement related to travel time before using this plan.
- Calculate the interval: .
- Randomize the start: use a uniform random generator to choose from . Suppose it is .
- Follow the interval: select positions , , , , …, .
- Check the count: the last position is , so the plan selects positions.
Method: systematic random sampling, because a random start is followed by a fixed interval.
Justification: the plan is easy to carry out on an ordered list and spreads selections along it. Its suitability depends on the list’s order. It is not an SRS of , because most possible sets of students cannot be formed by taking every position.
Do not start by choosing any number from and then simply stop at the end: that procedure may fail to select students and is not the stated plan.
Write a complete sampling plan in context
A reader should be able to follow your procedure without making a new selection decision. “Choose students randomly” identifies an intention, but leaves the actual method unclear.
Procedure frame: “Use [complete frame] for [defined population]. Label [units] from [range]. Use [random mechanism] to select [number] [distinct labels / groups / starting position]. [State the within-group or interval rule.] Collect [information] from [selected units].”
Justification frame: “This method fits because [a feature of this population or collection setting] makes [a specific design feature] useful. A limitation is [relevant tradeoff].”
Name evidence, not just a label
A complete identification sounds like: “This is stratified random sampling because the school divides students by grade and takes an SRS from each grade.” For cluster sampling, state that all students in selected groups are included. For systematic sampling, identify both the random start and the interval.
Keep the responding sample in view
The selected sample and the group that actually supplies data may differ. If a randomly selected student does not respond, inviting a convenient replacement is not automatically the same design. Plan an appropriate contact and follow-up process; distinguish missing data from a .
Use student IDs in the working selection record and collect only the information the question needs. The examples here require no private student information. Sampling labels identify selection units; they are not measurements of travel time.
Finally, keep the scope tied to the population covered by the frame. A complete Cedar High enrollment list supports selection for that school, not all teenagers. Sampling does not turn an observational association into a causal treatment effect.
Common mistakes and how to fix them
| Mistake | Why it fails | Better approach |
|---|---|---|
| “Everyone has an equal chance, so it must be an SRS.” | A design can restrict which groups of units are possible. | Check whether every possible sample of the required size has the same chance. |
| Calling a sample random because nobody deliberately chose the names. | An arbitrary or convenient choice need not use chance. | Identify the actual random mechanism. |
| Calling “some from every group” cluster sampling. | One-stage cluster sampling includes all units in selected groups. | For stratified random sampling, take an SRS from each stratum. |
| Calling every group-based plan stratified random. | The method depends on how groups and units are selected. | Distinguish samples from every group, whole selected groups, and multistage selection. |
| Keeping a repeated label in a sample of distinct units. | It counts the same unit twice under a no-replacement rule. | Skip the repeat and continue until the required distinct count is reached. |
| Starting every selection at position without randomizing. | The random-start feature is missing. | Choose uniformly among positions before following the interval. |
| Claiming stratified samples must use equal counts per group. | Allocation can be proportional or chosen for another justified goal. | State each stratum’s size and selection rule; consider weighting when fractions differ. |
| “Random sampling guarantees a representative sample and proves a cause.” | Chance variation remains; selection does not assign treatments. | Explain population coverage separately from experimental assignment. |
Eight practice questions with hints and solutions
These are original fictional scenarios. Give the method or calculation requested and explain it using the actual selection process.
1. Identify an SRS
A library has a complete list of registered members. Each receives a unique label . A uniform random generator selects distinct labels, and those members are surveyed. Identify the sampling method and replacement rule.
Hint
Are labels chosen from the whole frame? Can a label be counted twice?
Solution and explanation
SRS without replacement. The generator chooses distinct labels uniformly from the complete membership frame. Every set of members has the same selection chance, and no member is selected twice.
2. Equal individual chances, restricted samples
A fair coin selects either students A and B or students C and D. Each student has a chance of inclusion. Is this an SRS of students ?
Hint
Could A and C ever form the selected pair?
Solution and explanation
No. Only and can be selected. , , and have , while and each have chance . Equal individual inclusion chances do not give equal chances to all possible pairs.
3. Apply the replacement rule
From labels , uniform random draws give . Compare the interpretation with replacement and without replacement when distinct students are required.
Hint
Count the draws separately from the distinct labels.
Solution and explanation
With replacement: all draws are valid, including the repeated ; there are distinct students. Without replacement: accept , skip the repeated , and continue drawing until a valid label is accepted.
4. Allocate within strata
A school’s grade groups contain and students. It wants a proportional stratified random sample of . How many students should it select from each grade? Would selecting at random from each still be stratified random sampling?
Hint
Multiply each population share by . Separate the method from its allocation.
Solution and explanation
The population total is . The allocations are and . An SRS of from each grade would still be stratified random sampling, but it would not be proportional. It overrepresents the smaller grade in the unweighted sample, so whole-school estimation must account for the unequal population shares.
5. Select whole clusters
A camp has cabins, each with campers. Every camper belongs to . The director selects an SRS of cabins and surveys every camper in them. Identify the sampling method and planned number of sampled campers. Is it a census of the camp?
Hint
What is randomly selected, and who inside the selected groups is included?
Solution and explanation
One-stage cluster random sampling. Cabins are randomly selected clusters, and all campers in them are included. The planned sample is . It is not a census of the entire camp’s campers, even though every camper in the is surveyed.
6. Find the interval and endpoint
A complete list contains items. A systematic random sample will contain items. Determine the interval and random-start range. If the start is , give the and the last position.
Hint
Use , then add between selections. There are steps after the first selection.
Solution and explanation
. Choose the start uniformly from . Starting at gives , and initially. The last position is . This is a systematic random sample, not an SRS of , because the interval restricts possible samples.
7. Justify a method
A library system has branches. Each member is registered at , complete member lists are available, and borrowing patterns are expected to differ across branches. The researcher wants all branches to contribute to a member survey. Suggest and justify a random sampling method.
Hint
Which method samples individuals within every defined subgroup?
Solution and explanation
Use stratified random sampling by registered branch. Assign a positive sample size to each branch and take an SRS without replacement from each branch’s complete member list. This gives selected representation from all branches and may improve precision because borrowing patterns differ across branches. Specify allocation and account for population shares if using unequal selection fractions.
8. Read random digits with a stopping rule
are labeled . Select an SRS of distinct members by reading these successive groups: . Which labels are accepted, and where do you stop?
Hint
Skip invalid labels and repeats. Stop after distinct valid labels, not groups read.
Solution and explanation
Skip , accept , skip , skip the repeated , then accept , and . Stop at : the are . Do not include after the target is reached.
Quick revision checklist
Quick questions students often ask
Are strata always the same size?
No. Strata can have different population sizes, and their selected sample sizes can differ too. The defining feature is random selection within every stratum, not equal group sizes.
Is a class always a cluster?
A class can be used as a cluster when entire randomly selected classes are included. Taking an SRS of students from every class is stratified random sampling by class. Randomly selecting some classes and then sampling some students in them is multistage sampling. Read the actions, not just the group name.
Can systematic sampling be random without being an SRS?
Yes. Randomizing the start provides chance selection, but the fixed interval on an ordered list usually restricts the sets that can be selected. The makes the same distinction between individual chances and possible samples.
Does random sampling fix an incomplete population list?
No. A random procedure applied to an incomplete frame cannot select units missing from that frame. It also does not fix inaccurate answers or missing responses. Those issues are developed in Topic 1.12.
Final understanding check
A fictional school has enrolled students living in non-overlapping travel zones: Zone A has , Zone B has and Zone C has . Complete zone-specific student lists are available. Travel time is expected to differ across zones. The school wants to select students to estimate mean travel time on a specified Tuesday.
- Propose a sampling method and justify why it fits.
- Find the sample size for each zone under proportional allocation.
- Write a random selection procedure that produces distinct students.
- Explain why the combined sample is not an SRS of from all students.
- Assess a different plan: randomly select and survey every student living in it. Identify its method and a possible practical drawback.
Reveal the full solution
1. Method and reason: stratified random sampling by travel zone. Sampling from every zone gives selected representation for each area and can improve precision because travel times are expected to differ across zones.
2. Allocation: from A; from B; from C. The total is , and every zone’s selection fraction is .
3. Procedure: on each complete zone list, label students uniquely: , and . Independently within each zone, use a uniform random generator to select , or distinct valid numbered labels respectively. Prevent or ignore repeats within that zone, continuing until its target is reached. Combine the selected students and collect the same defined Tuesday travel-time response from each.
4. SRS distinction: the allocation always contains from A, from B and from C. Other sets of , such as a set entirely from Zone A, have . Not every possible sample of is equally likely.
5. Alternative: selecting randomly and including every student in it is one-stage cluster random sampling using zones as clusters. Its student count is , or , so it does not give the target . Because travel times are expected to differ between zones, collecting from only could give a sample poorly reflecting the full range of travel patterns. The random cluster choice does not guarantee an individual sample is representative.
Ready to move on? You should be able to identify who or what is randomized, explain replacement, distinguish strata from clusters, and write a selection plan someone else could carry out.
Continue learning
Topic 1.12: Potential Problems with Sampling →
Learn to identify undercoverage, nonresponse, voluntary response and response bias, and explain how they can affect a study.
Previous: Topic 1.10 — Investigative questions and data collection · Review the sampling methods · Review SRS and replacement · Back to the lesson overview