Summary Statistics for One Quantitative Variable
Give a distribution a useful numerical summary. Learn how to calculate and interpret center, position and spread, check possible outliers, and choose statistics that fit the data.
By the end of this lesson, you should be able to:
- Find and interpret the mean, median, quartiles, percentiles and extreme values.
- Calculate range, , sample standard deviation and sample variance with correct units.
- Apply two outlier rules and explain how a change of units affects summaries.
- Compare independent samples and justify a choice of center and spread.
Before you start: Know how to read a dotplot and describe shape. Review Topic 1.6 if you need a reminder.
First time learning this? Start with the Cedar High data, then follow the calculation steps in the worked examples.
Here to revise? Use the formula checklist, then try the practice questions before opening the solutions.
The concept in 60 seconds
A summary statistic turns a feature of recorded data into a number. For one quantitative variable, ask where the values are centered, how spread out they are, and where a value falls within the ordered data.
Center
Mean: add the values and divide by their number.
Median: find the middle of the ordered values.
Spread
Range: .
: .
Standard deviation: describes distances from the mean.
Position
Quartiles mark roughly the , and positions.
Percentiles describe relative standing within a group.
Numbers need context. “The median is ” describes the middle recorded travel time. “The is ” describes the width of the middle half. Those two numbers answer different questions.
Use a graph alongside the summaries. Two very different distributions can share a mean and standard deviation. A summary table alone does not reveal every cluster, gap or detail of shape.
From a travel-time graph to numbers
Use the same fictional sample of Cedar High students from Topics 1.5 and 1.6. Each value is a student’s one-way home-to-school travel time on the Tuesday being studied, rounded to the nearest minute.
Ordered travel times, minutes:
, , , , , , , , , ,
, , , , , , , , ,
dot represents student ·
On a small screen, scroll the plot sideways to read the full scale.
| Feature | Summary | What it tells us |
|---|---|---|
| Center | ; | The average is above the middle ordered value. |
| Position | ; | These mark the lower and upper quartile positions using our stated method. |
| Spread | ; ; | The full extent, middle-half width and variability around the mean are different measures. |
Predict: If the -minute journey became much longer, which center would usually move more: the mean or median?
Check your prediction
The mean would move more here because it uses the size of every value. As long as that journey remains the largest, changing it does not change the two middle observations, so this sample’s median stays .
The summaries describe these sampled students on the specified Tuesday. They do not by themselves establish a population pattern or explain why a journey took longer.
Mean, median, quartiles and percentiles
Mean: share the total equally
The sample mean, written (“x-bar”), is the total of the sample values divided by the sample size .
means “add”; is the value of ; is the number of observations.
For Cedar High, the total is , so . Imagine redistributing the total travel time equally among the observations: each would receive . The mean does not have to be an observed value.
Median: sort, then find the middle
- Order the values from smallest to largest.
- For odd , take the value at position .
- For even , average the values at positions and .
With values, the middle positions are and : and . The median is . It separates the ordered data into a lower half and an upper half; ties may prevent exactly half from being strictly below it.
Quartiles: find the middle of each half
Method used throughout these notes: find the median, then take the median of the lower half as and the median of the upper half as . For odd , leave the overall median out of both halves. Quartile conventions can differ between software and textbooks. Follow the method specified in your question or by your teacher.
· · · · · · · · ·
· · · · · · · · ·
The median is also . The minimum, , median, and maximum form a five-number summary: here , , , and . Topic 1.8 turns this information into a boxplot.
Percentiles: locate a value within a group
The percentile is a value at or below which approximately of the observations fall. The first quartile, median and third quartile correspond to approximately the , and percentiles.
When a question asks for the percentage at or below a specified value, count observations including ties at that value:
For Cedar High, of the travel times are at or below : . A travel time of has an inclusive empirical percentile rank of in this sample.
Read the wording carefully. For , the count at or below is , but the count strictly below is . Percentile rank describes standing within a group; it is different from a test score expressed as a percentage of available marks.
If you are asked to calculate a percentile cutoff
Use the stated rank or interpolation rule. For example, if a question specifies the nearest-rank method, use position in the ordered data, rounding the position upward. For and , position gives . This optional percentile method is separate from the median-of-halves quartile method used here; do not switch conventions silently.
With small samples or tied values, a computed percentile cutoff need not place exactly of observations at or below it. Different accepted methods may give different cutoffs.
Range, , sample standard deviation and variance
Range
Cedar High: . It measures the full numerical extent and uses only the two extremes.
Interquartile range
Cedar High: . It measures the width of the middle half of the ordered data.
Standard deviation: use distances from the mean
The sample standard deviation summarizes how spread out the observations are around . It is based on their squared deviations from the mean.
is the sample variance. Use to calculate these sample measures.
- Calculate the mean .
- For each observation, subtract the mean: .
- Square each deviation, then add the squared deviations.
- Divide by to get sample variance .
- Take the square root to get sample standard deviation .
Consider fictional homework completion times: , , , and . Their mean is .
homework times ·
On a small screen, scroll the plot sideways to read the full scale.
The signed deviations add to , so simply averaging them would hide the spread. Squaring prevents negative and positive deviations from cancelling. Taking the square root returns the standard deviation to the original units.
Interpretation: a sample standard deviation of about indicates a typical distance from the -minute sample mean, using the standard-deviation measure. It is not the arithmetic average of the absolute distances. It also does not mean that every observation is within of the mean.
Units matter
- Mean, median, quartiles, range, and standard deviation have the same units as the variable.
- Variance has squared units: , or .
- Standard deviation is never negative. It is when all observations are equal.
For the Cedar High sample, and . Keep extra digits during calculations and round the reported answer appropriately.
Why divide by ?
Once the sample mean is calculated, the deviations must add to . If deviations are known, the last is determined. This is the idea behind the degrees of freedom used in sample variance.
If the task instead asks for the standard deviation of an entire population, the population formula uses its population size in the denominator. Do not choose the population output merely because a calculator shows it next to the sample output.
Two outlier rules—and why they can disagree
An outlier rule flags an observation that is unusually far from the others according to a specified criterion. Apply the rule in the question; report which rule you used.
| Rule | Flag a value when… | Remember |
|---|---|---|
| The boundaries are called fences. A value exactly at a fence is not beyond it. | ||
| More than standard deviations from the mean | Use the stated mean and standard deviation; equality at a boundary is not more than away. |
Common minute scale · dashed lines show rule boundaries · a ring marks a flagged observation
On a small screen, scroll the plot sideways to read the full scale.
For Cedar High, is inside the fences but above . The two rules give different classifications. That is possible because one uses quartiles and the other uses the mean and standard deviation.
A flag is a reason to investigate. It does not prove a measurement error and does not automatically justify removing an observation. A negative lower travel-time fence is a calculated boundary; it does not mean a negative travel time was recorded.
Resistance: how strongly can an extreme value change a summary?
A resistant statistic is less affected by extreme values. The median and are resistant. The mean, range and standard deviation are not resistant; variance is also affected strongly by extreme values.
Same observations except the largest · common horizontal scale
On a small screen, scroll the plot sideways to read the full scale.
| Summary | Original values | Alternative values |
|---|---|---|
| Mean | ||
| Median | ||
| Range | ||
| Sample standard deviation |
Resistant does not mean “never changes.” The median and happen to stay the same in this example. Changing enough observations, or changing observations near the middle or quartile positions, can change them.
For a reasonably symmetric distribution without strong outliers, the mean and standard deviation are useful summaries. For a strongly skewed distribution or one with extreme outliers, the median and usually give a more stable description of center and spread. Still report additional graph features when the question asks for them.
Four worked examples
Example 1: summarize the Cedar High travel times
Prompt: Calculate the mean, median, quartiles, range and . Interpret the median and .
- Mean: and , so .
- Median: positions and contain and ; .
- Quartiles: lower-half middle values are and , giving . Upper-half middle values are and , giving .
- Range: . : .
Model interpretation: For these sampled students, the middle recorded one-way travel time is . The middle half of the ordered travel times extends from about to , an interval wide.
The is a width, rather than an endpoint, an average or a statement that every journey lasted .
Example 2: calculate sample standard deviation by hand
Prompt: Find the sample variance and standard deviation of the homework times , , , and .
The mean is . Build a table:
| (min) | (min) | |
|---|---|---|
The squared deviations total . Divide by : . Take the square root: .
Model interpretation: The sampled homework completion times have a mean of and a sample standard deviation of about , describing their variability around that mean.
Dividing by instead would give the population calculation, . It does not answer this question asking for sample standard deviation.
Example 3: check minutes using both rules
Prompt: Is Cedar High’s -minute journey flagged by the rule? By the more-than- rule?
- fences: ; . All observations, including , lie within these fences.
- boundaries: gives about and . Since , the -minute journey is flagged by this rule.
Model response: The -minute recorded journey is not flagged by the rule, because it is below the upper fence of . It is flagged by the more-than- rule, because it exceeds the upper boundary of about .
Use unrounded intermediate values for close boundary decisions. This screening calculation does not establish a normal distribution or a fixed percentage of observations outside the boundaries.
Example 4: change units or add a constant
Prompt: Convert the Cedar High summaries from minutes to seconds. Then consider a separate change that adds to every recorded time.
| Summary | Original, minutes | Seconds: | Add : |
|---|---|---|---|
| Mean | |||
| Median | |||
| , | , | , | , |
| Range | |||
Why? Multiplying every value by multiplies every distance by . Adding shifts the entire distribution without changing distances between observations.
For , with :
Mean, median and quartile values: .
Range, and : .
Variance: .
: use for range, and ; variance still uses . Order reverses, so the new comes from and the new from . The mean and median still follow .
Compare samples and explain what the numbers mean
A good response names the variable, group, statistic, numerical evidence and units. When comparing samples, use a direct comparison rather than two disconnected descriptions.
Compare independent samples
Two fictional offices record application processing times for independent samples of applications each. The observations are not paired. Using the stated quartile method:
| Summary | Office A | Office B |
|---|---|---|
| Mean and median | ; | ; |
| and | ; | ; |
| Sample standard deviation | ||
| Minimum and maximum | ; | ; |
Model comparison: Office B’s sampled processing times have a higher center: its median is , compared with for Office A, a difference of . They also show greater variability: B’s is , twice A’s , and B’s sample standard deviation is twice A’s.
This does not imply every B time is larger than every A time: B’s minimum is , below A’s . The table alone also does not show the exact distribution shape or prove an office caused a delay. A claim about all applications would require information about the sampling process.
See the invented values behind the comparison
A: , , , , , , , , .
B: , , , , , , , , .
Justify your choice of statistics
Connect the choice to a feature of the distribution. For the alternative service times with an isolated -minute value, write: “I would report the median of and of because the distribution has a high extreme value, and these summaries are resistant to its effect.”
“The median is better” is incomplete. Explain the feature, the resistance, and the numbers. If two clusters are important, describe them as well; one center and spread pair can conceal the groups.
Calculator workflow: check the label and the data
- Enter all observations into a list. On a TI-84 family calculator, use STAT → EDIT and enter the values in L1. Clear any old entries first.
- Run 1-Var Stats from STAT → CALC, using L1. Leave the frequency list blank for a simple list containing each observation once.
- Check and first. For the homework times, should be and should be .
- Read for sample standard deviation. For these times, . The separate output, about , uses the population denominator.
- Square for sample variance if needed. Check , , , and , remembering the quartile convention.
Menu wording varies by model. If using a frequency list, enter each distinct value with its count and check that equals the total count, rather than the number of distinct values. Keep calculator precision until the final reported answer.
AP response habit: show the setup or identify the statistic, then give the result and interpret it. Calculator output by itself does not explain what a number means in the study.
Find and fix common mistakes
| Mistake | Why it fails | Better approach |
|---|---|---|
| Finding the median before sorting | The middle entry in an unsorted list is not necessarily the median. | Order the observations and count positions. |
| Calling the Cedar High range | is the maximum, not a difference. | . |
| Using in a sample-variance calculation | It gives the population-denominator result. | Use for and read the sample calculator output . |
| Reporting variance in minutes | Squared deviations have squared units. | Report variance in and standard deviation in minutes. |
| Saying “ percentile means correct” | Relative standing and percentage marks answer different questions. | Identify the comparison group and the proportion at or below the value. |
| Flagging a point exactly on an fence | The rule requires a point beyond the fence. | Use . |
| Adding to a standard deviation when converting to | An additive shift does not change distances. | Multiply by ; multiply variance by . |
| Deleting every flagged outlier | A flag does not establish an error. | Investigate the observation and justify any data-handling decision. |
| Treating as proof of right skew | That comparison can suggest skew but does not determine the full shape. | Check a graph and support shape claims with the pattern. |
Eight practice questions with hints and solutions
Write a calculation and a sentence with units when the question asks for interpretation. All datasets below are invented teaching examples.
1. Center, quartiles and spread
students report reading times of , , , , , and . Find the mean, median, , , range and . Use our median-of-halves method, excluding the overall median for odd .
Need a hint?
The total is . The value is the median. Split the remaining values into groups of .
Show the solution
; . Lower half , , gives ; upper half , , gives . ; .
The mean and middle recorded reading time are both , and the middle half spans an interval wide.
2. Sample or population output?
A sample of takes , and . Find and . A calculator shows and . Which output answers the question?
Need a hint?
Find the mean, square deviations of , and , and divide their total by .
Show the solution
. The squared deviations total . ; . Use . The output divides by and corresponds to the population calculation.
3. Percentile wording and ties
students score , , , , , , , , and points. What percentage scored at or below ? What percentage scored strictly below ? Explain why those answers differ.
Need a hint?
Include both -point scores for “at or below”; exclude both for “strictly below.”
Show the solution
At or below : . Strictly below : . The tied scores account for the difference. Using the inclusive empirical convention, has percentile rank within this group. That describes standing among students, rather than the percentage of possible test marks earned.
4. Exactly on an outlier fence
A temperature dataset has and . Among readings of , and , which are flagged by the rule?
Need a hint?
. Find both fences and apply strict inequalities.
Show the solution
. . Only is beyond a fence. and are exactly on the fences and are not flagged by this rule.
5. Explain a disagreement between rules
For the Cedar High travel times, , , and , all in minutes. Check the -minute journey using each rule. Does a flag prove a recording error?
Need a hint?
Compare with and .
Show the solution
; . Since is below that fence, it is not flagged by this rule. The upper boundary is about , so it is flagged by the more-than- rule. State the rule with the result. The flag does not prove an error or justify automatic deletion.
6. Convert temperature summaries
A sample has mean , median , , and . Convert these summaries to Fahrenheit using .
Need a hint?
Apply the full transformation to centers. Apply only the multiplier to spread, and square the multiplier for variance.
Show the solution
; . ; ; . Adding changes location but not spread.
7. Compare without overclaiming
Office A’s sampled processing times have median and . Office B’s have median and . Compare their center and spread. Can you conclude every B application took longer, or that all B applications take longer?
Need a hint?
Use “higher” and “greater variability” with numerical evidence. Separate sample summaries from claims about every observation or a population.
Show the solution
B’s sampled processing times have a -minute higher median and an twice as large as A’s. The middle half of B’s times therefore covers a wider interval. These summaries do not imply every B observation is larger. Generalizing to all applications also needs information about how the samples were obtained.
8. Choose summaries and justify
For the service-time sensitivity example, changing the largest value from to moves the mean from to and from to . The median stays and stays . Which center and spread would you report for the alternative distribution, and why?
Need a hint?
Look at the isolated high value and connect it to resistance.
Show the solution
I would report the median of and of because the alternative service-time distribution has an isolated high value, and these summaries are resistant to its effect. The changes in the mean and standard deviation illustrate their sensitivity. This example does not mean resistant summaries can never change.
Quick revision checklist
| Task | Tool | Check before reporting |
|---|---|---|
| Find an average | Total, sample size and units. | |
| Find the middle and quartiles | Sort; locate middle positions and half medians. | Odd/even , ties and stated quartile convention. |
| Describe relative standing | Percentile or percentage at/below a cutoff. | Reference group and . |
| Find spread | Range; ; ; . | Spread units; variance units squared; sample output. |
| Check possible outliers | ; or . | Specified rule, strict inequality and rounding. |
| Change units: , | Centers: ; spreads: ; variance: . | . |
| Choose a center/spread pair | Mean with for reasonably symmetric data without strong outliers; median with for strong skew or outliers. | Support the choice with a distribution feature. |
| Compare independent samples | Directly compare the same statistic, with differences or ratios. | Variable, group, units and limits of the summaries. |
Before submitting: Have I identified the statistic, used the right units, shown enough calculation, answered in context and avoided a stronger claim than the data support?
Final understanding check
sampled visitors have recorded waits of , , , , , , and .
dot represents sampled visitor ·
On a small screen, scroll the plot sideways to read the full scale.
- Calculate the mean, median, quartiles, range and using our stated method.
- The squared deviations from the mean total . Find and .
- Apply the rule. Choose a center/spread pair and justify your choice.
- What percentage of waits is at or below ? Convert the median and to seconds.
Show the full check and explanation
1. The total is , so . . Lower half , , , gives ; upper half , , , gives . ; .
2. ; .
3. and . The -minute wait exceeds the upper fence and is flagged. The median of and of are useful because the distribution has an isolated high value and these summaries are resistant.
4. are at or below : . Multiplying by gives a median of and of .
Model contextual summary: For these sampled visitors, the middle recorded wait is and the middle half spans about . The -minute wait is flagged by the rule. These statements describe this recorded sample; they do not identify the reason for the long wait.
Ready to move on? You should be able to calculate a summary, interpret it with units, apply a stated outlier rule, transform units and justify a comparison or a choice of statistics.
Continue learning
Use quartiles and outlier fences to build and interpret boxplots, then compare distributions visually.
Review quartiles and percentiles · Review spread formulas · Back to the lesson overview