AP Statistics  /  Unit 1: Exploring One-Variable Data and Collecting Data  /  Topic 1.7
NUM8ERS study notes · Topic 1.7

Summary Statistics for One Quantitative Variable

Give a distribution a useful numerical summary. Learn how to calculate and interpret center, position and spread, check possible outliers, and choose statistics that fit the data.

2026–27 curriculum4 worked examples8 practice questionsCalculate → interpret → justify

By the end of this lesson, you should be able to:

  • Find and interpret the mean, median, quartiles, percentiles and extreme values.
  • Calculate range, IQR⁡\operatorname{IQR}, sample standard deviation and sample variance with correct units.
  • Apply two outlier rules and explain how a change of units affects summaries.
  • Compare independent samples and justify a choice of center and spread.

Before you start: Know how to read a dotplot and describe shape. Review Topic 1.6 if you need a reminder.

First time learning this? Start with the Cedar High data, then follow the calculation steps in the worked examples.

Here to revise? Use the formula checklist, then try the practice questions before opening the solutions.

The concept in 60 seconds

A summary statistic turns a feature of recorded data into a number. For one quantitative variable, ask where the values are centered, how spread out they are, and where a value falls within the ordered data.

Center

Mean: add the values and divide by their number.
Median: find the middle of the ordered values.

Spread

Range: maximum−minimum\text{maximum}-\text{minimum}.
IQR⁡\operatorname{IQR}: Q3−Q1Q_3-Q_1.
Standard deviation: describes distances from the mean.

Position

Quartiles mark roughly the 25%25\%, 50%50\% and 75%75\% positions.
Percentiles describe relative standing within a group.

Numbers need context. “The median is 19 min19 \,\mathrm{min}” describes the middle recorded travel time. “The IQR⁡\operatorname{IQR} is 12.5 min12.5 \,\mathrm{min}” describes the width of the middle half. Those two numbers answer different questions.

Use a graph alongside the summaries. Two very different distributions can share a mean and standard deviation. A summary table alone does not reveal every cluster, gap or detail of shape.

From a travel-time graph to numbers

Use the same fictional sample of 2020 Cedar High students from Topics 1.5 and 1.6. Each value is a student’s one-way home-to-school travel time on the Tuesday being studied, rounded to the nearest minute.

Ordered travel times, minutes:
88, 1010, 1212, 1414, 1414, 1515, 1616, 1818, 1818, 1818,
2020, 2222, 2424, 2525, 2626, 2828, 3030, 3232, 3535, 4242

Mean and median locate the same data differently

11 dot represents 11 student · n=20n = 20

Cedar High travel times with mean and medianTwenty travel times: 8, 10, 12, 14, 14, 15, 16, 18, 18, 18, 20, 22, 24, 25, 26, 28, 30, 32, 35, 42. Dashed line at median 19 minutes; dotted line at mean 21.35 minutes. 5 10 15 20 25 30 35 40 45 Recorded one-way travel time (minutes) Median = 19 min Mean = 21.35 min

On a small screen, scroll the plot sideways to read the full scale.

The longer travel times pull the mean above the median. Neither center has to be an observed travel time. Invented teaching data.
Numerical summaries of the 2020 recorded travel times.
FeatureSummaryWhat it tells us
CenterMean=21.35 min\text{Mean} = 21.35 \,\mathrm{min}; median=19 min\text{median} = 19 \,\mathrm{min}The average is above the middle ordered value.
PositionQ1=14.5 minQ_{1} = 14.5 \,\mathrm{min}; Q3=27 minQ_{3} = 27 \,\mathrm{min}These mark the lower and upper quartile positions using our stated method.
SpreadRange=34 min\text{Range} = 34 \,\mathrm{min}; IQR⁡=12.5 min\operatorname{IQR} = 12.5 \,\mathrm{min}; s≈8.87 mins \approx 8.87 \,\mathrm{min}The full extent, middle-half width and variability around the mean are different measures.

Predict: If the 4242-minute journey became much longer, which center would usually move more: the mean or median?

Check your prediction

The mean would move more here because it uses the size of every value. As long as that journey remains the largest, changing it does not change the two middle observations, so this sample’s median stays 19 min19 \,\mathrm{min}.

The summaries describe these 2020 sampled students on the specified Tuesday. They do not by themselves establish a population pattern or explain why a journey took longer.

Mean, median, quartiles and percentiles

Mean: share the total equally

The sample mean, written xˉ\bar{x} (“x-bar”), is the total of the sample values divided by the sample size nn.

xˉ=∑i=1nxin\bar{x}=\frac{\sum_{i=1}^{n}x_i}{n}

∑\sum means “add”; xix_i is the value of observation i\text{observation }i; nn is the number of observations.

For Cedar High, the total is 427 min427 \,\mathrm{min}, so xˉ=42720=21.35 min\bar{x}=\frac{427}{20}=21.35\,\mathrm{min}. Imagine redistributing the total travel time equally among the 2020 observations: each would receive 21.35 min21.35 \,\mathrm{min}. The mean does not have to be an observed value.

Median: sort, then find the middle

  1. Order the values from smallest to largest.
  2. For odd nn, take the value at position n+12\frac{n+1}{2}.
  3. For even nn, average the values at positions n2\frac{n}{2} and n2+1\frac{n}{2}+1.

With 2020 values, the middle positions are 1010 and 1111: 1818 and 20 min20 \,\mathrm{min}. The median is 18+202=19 min\frac{18+20}{2}=19\,\mathrm{min}. It separates the ordered data into a lower half and an upper half; ties may prevent exactly half from being strictly below it.

Quartiles: find the middle of each half

Method used throughout these notes: find the median, then take the median of the lower half as Q1Q_{1} and the median of the upper half as Q3Q_{3}. For odd nn, leave the overall median out of both halves. Quartile conventions can differ between software and textbooks. Follow the method specified in your question or by your teacher.

Lower half · positions 1–101\text{–}10

88 · 1010 · 1212 · 1414 · 1414 · 1515 · 1616 · 1818 · 1818 · 1818

Q1=14+152=14.5 minQ_1=\frac{14+15}{2}=14.5\,\mathrm{min}

Upper half · positions 11–2011\text{–}20

2020 · 2222 · 2424 · 2525 · 2626 · 2828 · 3030 · 3232 · 3535 · 4242

Q3=26+282=27 minQ_3=\frac{26+28}{2}=27\,\mathrm{min}

The median is also Q2Q_{2}. The minimum, Q1Q_{1}, median, Q3Q_{3} and maximum form a five-number summary: here 88, 14.514.5, 1919, 2727 and 42 min42 \,\mathrm{min}. Topic 1.8 turns this information into a boxplot.

Percentiles: locate a value within a group

The pthp^{\mathrm{th}} percentile is a value at or below which approximately p%p\% of the observations fall. The first quartile, median and third quartile correspond to approximately the 25th25^{\mathrm{th}}, 50th50^{\mathrm{th}} and 75th75^{\mathrm{th}} percentiles.

When a question asks for the percentage at or below a specified value, count observations including ties at that value:

Percentage at or below a cutoff=count at or below the cutoffn×100%\text{Percentage at or below a cutoff}=\frac{\text{count at or below the cutoff}}{n}\times100\%

For Cedar High, 1616 of the 2020 travel times are at or below 28 min28 \,\mathrm{min}: 1620×100%=80%\frac{16}{20}\times100\%=80\%. A travel time of 28 min28 \,\mathrm{min} has an inclusive empirical percentile rank of 8080 in this sample.

Read the wording carefully. For 25 min25 \,\mathrm{min}, the count at or below is 1420=70%\frac{14}{20}=70\%, but the count strictly below is 1320=65%\frac{13}{20}=65\%. Percentile rank describes standing within a group; it is different from a test score expressed as a percentage of available marks.

If you are asked to calculate a percentile cutoff

Use the stated rank or interpolation rule. For example, if a question specifies the nearest-rank method, use position ⌈pn100⌉\left\lceil\frac{pn}{100}\right\rceil in the ordered data, rounding the position upward. For p=80p = 80 and n=20n = 20, position 1616 gives 28 min28 \,\mathrm{min}. This optional percentile method is separate from the median-of-halves quartile method used here; do not switch conventions silently.

With small samples or tied values, a computed percentile cutoff need not place exactly p%p\% of observations at or below it. Different accepted methods may give different cutoffs.

Range, IQR⁡\operatorname{IQR}, sample standard deviation and variance

Range

Maximum−minimum\text{Maximum}-\text{minimum}

Cedar High: 42−8=34 min42-8=34\,\mathrm{min}. It measures the full numerical extent and uses only the two extremes.

Interquartile range

IQR⁡=Q3−Q1\operatorname{IQR}=Q_3-Q_1

Cedar High: 27−14.5=12.5 min27-14.5=12.5\,\mathrm{min}. It measures the width of the middle half of the ordered data.

Standard deviation: use distances from the mean

The sample standard deviation ss summarizes how spread out the observations are around xˉ\bar{x}. It is based on their squared deviations from the mean.

s2=∑i=1n(xi−xˉ)2n−1s^{2}=\frac{\sum_{i=1}^{n}(x_i-\bar{x})^{2}}{n-1}

s=∑i=1n(xi−xˉ)2n−1s=\sqrt{\frac{\sum_{i=1}^{n}(x_i-\bar{x})^{2}}{n-1}}

s2s^{2} is the sample variance. Use n≥2n\ge2 to calculate these sample measures.

  1. Calculate the mean xˉ\bar{x}.
  2. For each observation, subtract the mean: xi−xˉx_i-\bar{x}.
  3. Square each deviation, then add the squared deviations.
  4. Divide by n−1n-1 to get sample variance s2s^{2}.
  5. Take the square root to get sample standard deviation ss.

Consider 55 fictional homework completion times: 1212, 1414, 1616, 1818 and 20 min20 \,\mathrm{min}. Their mean is 16 min16 \,\mathrm{min}.

Distances from the mean: below, at and above

55 homework times · mean=16 min\text{mean} = 16 \,\mathrm{min}

Signed deviations from the homework-time meanTimes 12,14,16,18,20 minutes have respective signed deviations -4,-2,0,2,4 minutes from mean 16. The signed deviations sum to zero. −4 −2 0 2 4 Signed deviation from the 16-minute mean (minutes) 12 min 14 min 16 min 18 min 20 min -4 -2 0 +2 +4

On a small screen, scroll the plot sideways to read the full scale.

The sign shows which side of the mean a value falls on. Squared deviations of 1616, 44, 00, 44 and 1616 total 40 min240 \,\mathrm{min}^{2}. Invented teaching data.

The signed deviations add to 00, so simply averaging them would hide the spread. Squaring prevents negative and positive deviations from cancelling. Taking the square root returns the standard deviation to the original units.

Interpretation: a sample standard deviation of about 3.16 min3.16 \,\mathrm{min} indicates a typical distance from the 1616-minute sample mean, using the standard-deviation measure. It is not the arithmetic average of the absolute distances. It also does not mean that every observation is within 3.16 min3.16 \,\mathrm{min} of the mean.

Units matter

  • Mean, median, quartiles, range, IQR⁡\operatorname{IQR} and standard deviation have the same units as the variable.
  • Variance has squared units:  min2\,\mathrm{min}^{2},  dollar2\,\mathrm{dollar}^{2} or  point2\,\mathrm{point}^{2}.
  • Standard deviation is never negative. It is 00 when all observations are equal.

For the Cedar High sample, s2≈78.66 min2s^{2} \approx 78.66 \,\mathrm{min}^{2} and s≈8.87 mins \approx 8.87 \,\mathrm{min}. Keep extra digits during calculations and round the reported answer appropriately.

Why divide by n−1n-1?

Once the sample mean is calculated, the deviations must add to 00. If n−1n-1 deviations are known, the last is determined. This is the idea behind the n−1n-1 degrees of freedom used in sample variance.

If the task instead asks for the standard deviation of an entire population, the population formula uses its population size in the denominator. Do not choose the population output merely because a calculator shows it next to the sample output.

Two outlier rules—and why they can disagree

An outlier rule flags an observation that is unusually far from the others according to a specified criterion. Apply the rule in the question; report which rule you used.

Two numerical screening rules.
RuleFlag a value when…Remember
1.5×IQR⁡1.5\times\operatorname{IQR}x<Q1−1.5IQR⁡orx>Q3+1.5IQR⁡x<Q_1-1.5\operatorname{IQR}\quad\text{or}\quad x>Q_3+1.5\operatorname{IQR}The boundaries are called fences. A value exactly at a fence is not beyond it.
More than 22 standard deviations from the meanx<xˉ−2sorx>xˉ+2sx<\bar{x}-2s\quad\text{or}\quad x>\bar{x}+2sUse the stated mean and standard deviation; equality at a boundary is not more than 2s2s away.
The same observations, two different screening rules

Common minute scale · dashed lines show rule boundaries · a ring marks a flagged observation

Cedar High outlier screening with two rulesSame twenty observations in both panels. IQR fences -4.25 and 45.75 minutes flag none. Mean plus or minus two sample standard deviations gives boundaries about 3.61 and 39.09 minutes, flagging only 42. A ring and arrow identify the flagged point. −10 0 10 20 30 40 50 Recorded one-way travel time (minutes) 1.5 × IQR rule · no observations flagged -4.25 45.75 −10 0 10 20 30 40 50 Recorded one-way travel time (minutes) 42 More-than-2s rule · 42 minutes flagged 3.61 39.09

On a small screen, scroll the plot sideways to read the full scale.

The 4242-minute journey lies below the upper IQR⁡\operatorname{IQR} fence but above the upper 2s2s boundary. Negative scale values accommodate calculated fences, rather than recorded negative journeys. Invented teaching data.

For Cedar High, 42 min42 \,\mathrm{min} is inside the IQR⁡\operatorname{IQR} fences but above xˉ+2s\bar{x}+2s. The two rules give different classifications. That is possible because one uses quartiles and the other uses the mean and standard deviation.

A flag is a reason to investigate. It does not prove a measurement error and does not automatically justify removing an observation. A negative lower travel-time fence is a calculated boundary; it does not mean a negative travel time was recorded.

Resistance: how strongly can an extreme value change a summary?

A resistant statistic is less affected by extreme values. The median and IQR⁡\operatorname{IQR} are resistant. The mean, range and standard deviation are not resistant; variance is also affected strongly by extreme values.

An extreme value changes some summaries more than others

Same 99 observations except the largest · common horizontal scale

Service-time sensitivity comparisonOriginal times 4,5,6,6,7,8,9,10,11. Alternative times 4,5,6,6,7,8,9,10,25. Both medians are 7 and both IQRs are 4. Mean changes from 7.33 to 8.89 and sample standard deviation from 2.35 to 6.33 minutes. 0 5 10 15 20 25 Recorded service time (minutes) 11 Original: largest service time is 11 minutes 0 5 10 15 20 25 Recorded service time (minutes) 25 Alternative: only the largest time becomes 25 minutes

On a small screen, scroll the plot sideways to read the full scale.

Compare the table below: the middle and quartile positions stay fixed in this particular example, while the total and squared deviations increase. Invented teaching data.
A fictional sensitivity comparison: change only the largest service time, from 1111 to 25 min25 \,\mathrm{min}.
SummaryOriginal 99 valuesAlternative 99 values
Mean7.33 min7.33 \,\mathrm{min}8.89 min8.89 \,\mathrm{min}
Median7 min7 \,\mathrm{min}7 min7 \,\mathrm{min}
IQR⁡\operatorname{IQR}4 min4 \,\mathrm{min}4 min4 \,\mathrm{min}
Range7 min7 \,\mathrm{min}21 min21 \,\mathrm{min}
Sample standard deviation2.35 min2.35 \,\mathrm{min}6.33 min6.33 \,\mathrm{min}

Resistant does not mean “never changes.” The median and IQR⁡\operatorname{IQR} happen to stay the same in this example. Changing enough observations, or changing observations near the middle or quartile positions, can change them.

For a reasonably symmetric distribution without strong outliers, the mean and standard deviation are useful summaries. For a strongly skewed distribution or one with extreme outliers, the median and IQR⁡\operatorname{IQR} usually give a more stable description of center and spread. Still report additional graph features when the question asks for them.

Four worked examples

Example 1: summarize the Cedar High travel times

Prompt: Calculate the mean, median, quartiles, range and IQR⁡\operatorname{IQR}. Interpret the median and IQR⁡\operatorname{IQR}.

  1. Mean: ∑i=1nxi=427\sum_{i=1}^{n}x_i=427 and n=20n = 20, so xˉ=42720=21.35 min\bar{x}=\frac{427}{20}=21.35\,\mathrm{min}.
  2. Median: positions 1010 and 1111 contain 1818 and 2020; median=19 min\text{median} = 19 \,\mathrm{min}.
  3. Quartiles: lower-half middle values are 1414 and 1515, giving Q1=14.5Q_{1} = 14.5. Upper-half middle values are 2626 and 2828, giving Q3=27Q_{3} = 27.
  4. Range: 42−8=34 min42-8=34\,\mathrm{min}. IQR⁡\operatorname{IQR}: 27−14.5=12.5 min27-14.5=12.5\,\mathrm{min}.

Model interpretation: For these 2020 sampled students, the middle recorded one-way travel time is 19 min19 \,\mathrm{min}. The middle half of the ordered travel times extends from about 14.514.5 to 27 min27 \,\mathrm{min}, an interval 12.5 min12.5 \,\mathrm{min} wide.

The IQR⁡\operatorname{IQR} is a width, rather than an endpoint, an average or a statement that every journey lasted 12.5 min12.5 \,\mathrm{min}.

Example 2: calculate sample standard deviation by hand

Prompt: Find the sample variance and standard deviation of the 55 homework times 1212, 1414, 1616, 1818 and 20 min20 \,\mathrm{min}.

The mean is 805=16 min\frac{80}{5}=16\,\mathrm{min}. Build a table:

Deviations from the 1616-minute mean.
xix_i (min)xi−xˉx_i-\bar{x} (min)(xi−xˉ)2  (min2)(x_i-\bar{x})^{2}\;(\mathrm{min}^{2})
1212−4-41616
1414−2-244
16160000
18182244
2020441616

The squared deviations total 40 min240 \,\mathrm{min}^{2}. Divide by n−1=4n-1=4: s2=404=10 min2s^{2}=\frac{40}{4}=10\,\mathrm{min}^{2}. Take the square root: s=10≈3.16 mins=\sqrt{10}\approx3.16\,\mathrm{min}.

Model interpretation: The 55 sampled homework completion times have a mean of 16 min16 \,\mathrm{min} and a sample standard deviation of about 3.16 min3.16 \,\mathrm{min}, describing their variability around that mean.

Dividing by 55 instead would give the population calculation, 8≈2.83 min\sqrt{8}\approx2.83\,\mathrm{min}. It does not answer this question asking for sample standard deviation.

Example 3: check 4242 minutes using both rules

Prompt: Is Cedar High’s 4242-minute journey flagged by the 1.5×IQR⁡1.5\times\operatorname{IQR} rule? By the more-than-2s2s rule?

  1. IQR⁡\operatorname{IQR} fences: lower=14.5−1.5(12.5)=−4.25\text{lower}=14.5-1.5(12.5)=-4.25; upper=27+1.5(12.5)=45.75 min\text{upper}=27+1.5(12.5)=45.75\,\mathrm{min}. All observations, including 4242, lie within these fences.
  2. 2s2s boundaries: 21.35±2(8.86907697…)21.35\pm2(8.86907697\ldots) gives about 3.613.61 and 39.09 min39.09 \,\mathrm{min}. Since 42>39.0942>39.09, the 4242-minute journey is flagged by this rule.

Model response: The 4242-minute recorded journey is not flagged by the 1.5×IQR⁡1.5\times\operatorname{IQR} rule, because it is below the upper fence of 45.75 min45.75 \,\mathrm{min}. It is flagged by the more-than-2s2s rule, because it exceeds the upper boundary of about 39.09 min39.09 \,\mathrm{min}.

Use unrounded intermediate values for close boundary decisions. This screening calculation does not establish a normal distribution or a fixed percentage of observations outside the boundaries.

Example 4: change units or add a constant

Prompt: Convert the Cedar High summaries from minutes to seconds. Then consider a separate change that adds 5 min5 \,\mathrm{min} to every recorded time.

Apply the transformation to summaries, rather than recalculating every observation.
SummaryOriginal, minutesSeconds: y=60xy=60xAdd 5 min5 \,\mathrm{min}: y=x+5y=x+5
Mean21.3521.351,281 s1{,}281 \,\mathrm{s}26.35 min26.35 \,\mathrm{min}
Median19191,140 s1{,}140 \,\mathrm{s}24 min24 \,\mathrm{min}
Q1Q_{1}, Q3Q_{3}14.514.5, 2727870870, 1,620 s1{,}620 \,\mathrm{s}19.519.5, 32 min32 \,\mathrm{min}
Range34342,040 s2{,}040 \,\mathrm{s}34 min34 \,\mathrm{min}
IQR⁡\operatorname{IQR}12.512.5750 s750 \,\mathrm{s}12.5 min12.5 \,\mathrm{min}
ss≈8.8691\approx 8.8691≈532.14 s\approx 532.14 \,\mathrm{s}≈8.87 min\approx 8.87 \,\mathrm{min}

Why? Multiplying every value by 6060 multiplies every distance by 6060. Adding 55 shifts the entire distribution without changing distances between observations.

For y=a+bxy=a+bx, with b>0b>0:

Mean, median and quartile values: a+b×original summarya+b\times\text{original summary}.

Range, IQR⁡\operatorname{IQR} and ss: b×original spreadb\times\text{original spread}.

Variance: b2×original varianceb^{2}\times\text{original variance}.

If b is negative\text{If }b\text{ is negative}: use ∣b∣|b| for range, IQR⁡\operatorname{IQR} and ss; variance still uses b2b^{2}. Order reverses, so the new Q1Q_{1} comes from a+bQ3a+bQ_3 and the new Q3Q_{3} from a+bQ1a+bQ_1. The mean and median still follow a+b×their original valuesa+b\times\text{their original values}.

Compare samples and explain what the numbers mean

A good response names the variable, group, statistic, numerical evidence and units. When comparing samples, use a direct comparison rather than two disconnected descriptions.

Compare independent samples

Two fictional offices record application processing times for independent samples of 99 applications each. The observations are not paired. Using the stated quartile method:

Processing-time summaries for the two recorded samples.
SummaryOffice AOffice B
Mean and median10 min10 \,\mathrm{min}; 10 min10 \,\mathrm{min}13 min13 \,\mathrm{min}; 13 min13 \,\mathrm{min}
Q1Q_{1} and Q3Q_{3}8 min8 \,\mathrm{min}; 12 min12 \,\mathrm{min}9 min9 \,\mathrm{min}; 17 min17 \,\mathrm{min}
IQR⁡\operatorname{IQR}4 min4 \,\mathrm{min}8 min8 \,\mathrm{min}
Sample standard deviation2.5 min2.5 \,\mathrm{min}5 min5 \,\mathrm{min}
Minimum and maximum6 min6 \,\mathrm{min}; 14 min14 \,\mathrm{min}5 min5 \,\mathrm{min}; 21 min21 \,\mathrm{min}

Model comparison: Office B’s sampled processing times have a higher center: its median is 13 min13 \,\mathrm{min}, compared with 10 min10 \,\mathrm{min} for Office A, a difference of 3 min3 \,\mathrm{min}. They also show greater variability: B’s IQR⁡\operatorname{IQR} is 8 min8 \,\mathrm{min}, twice A’s 4 min4 \,\mathrm{min}, and B’s sample standard deviation is twice A’s.

This does not imply every B time is larger than every A time: B’s minimum is 5 min5 \,\mathrm{min}, below A’s 66. The table alone also does not show the exact distribution shape or prove an office caused a delay. A claim about all applications would require information about the sampling process.

See the invented values behind the comparison

A: 66, 88, 88, 99, 1010, 1111, 1212, 1212, 14 min14 \,\mathrm{min}.
B: 55, 99, 99, 1111, 1313, 1515, 1717, 1717, 21 min21 \,\mathrm{min}.

Justify your choice of statistics

Connect the choice to a feature of the distribution. For the alternative service times with an isolated 2525-minute value, write: “I would report the median of 7 min7 \,\mathrm{min} and IQR⁡\operatorname{IQR} of 4 min4 \,\mathrm{min} because the distribution has a high extreme value, and these summaries are resistant to its effect.”

“The median is better” is incomplete. Explain the feature, the resistance, and the numbers. If two clusters are important, describe them as well; one center and spread pair can conceal the groups.

Calculator workflow: check the label and the data

  1. Enter all observations into a list. On a TI-84 family calculator, use STAT → EDIT and enter the values in L1. Clear any old entries first.
  2. Run 1-Var Stats from STAT → CALC, using L1. Leave the frequency list blank for a simple list containing each observation once.
  3. Check nn and xˉ\bar{x} first. For the 55 homework times, nn should be 55 and xˉ\bar{x} should be 1616.
  4. Read SxS_x for sample standard deviation. For these times, Sx≈3.1623S_x \approx 3.1623. The separate σx\sigma_x output, about 2.82842.8284, uses the population denominator.
  5. Square SxS_x for sample variance if needed. Check minX\mathtt{minX}, Q1Q_1, Med\mathtt{Med}, Q3Q_3 and maxX\mathtt{maxX}, remembering the quartile convention.

Menu wording varies by model. If using a frequency list, enter each distinct value with its count and check that nn equals the total count, rather than the number of distinct values. Keep calculator precision until the final reported answer.

AP response habit: show the setup or identify the statistic, then give the result and interpret it. Calculator output by itself does not explain what a number means in the study.

Find and fix common mistakes

Errors that change the calculation or weaken the interpretation.
MistakeWhy it failsBetter approach
Finding the median before sortingThe middle entry in an unsorted list is not necessarily the median.Order the observations and count positions.
Calling 4242 the Cedar High range4242 is the maximum, not a difference.Range=42−8=34 min\text{Range}=42-8=34\,\mathrm{min}.
Using nn in a sample-variance calculationIt gives the population-denominator result.Use n−1n-1 for s2s^{2} and read the sample calculator output SxS_x.
Reporting variance in minutesSquared deviations have squared units.Report variance in  min2\,\mathrm{min}^{2} and standard deviation in minutes.
Saying “80th80^{\mathrm{th}} percentile means 80%80\% correct”Relative standing and percentage marks answer different questions.Identify the comparison group and the proportion at or below the value.
Flagging a point exactly on an IQR⁡\operatorname{IQR} fenceThe rule requires a point beyond the fence.Use strict < or > inequalities\text{strict }<\text{ or }>\text{ inequalities}.
Adding 3232 to a standard deviation when converting ∘C{}^{\circ}\mathrm{C} to ∘F{}^{\circ}\mathrm{F}An additive shift does not change distances.Multiply ss by 1.81.8; multiply variance by 1.821.8^{2}.
Deleting every flagged outlierA flag does not establish an error.Investigate the observation and justify any data-handling decision.
Treating mean>median\text{mean}>\text{median} as proof of right skewThat comparison can suggest skew but does not determine the full shape.Check a graph and support shape claims with the pattern.

Eight practice questions with hints and solutions

Write a calculation and a sentence with units when the question asks for interpretation. All datasets below are invented teaching examples.

1. Center, quartiles and spread

77 students report reading times of 1212, 1515, 1818, 2020, 2020, 2424 and 31 min31 \,\mathrm{min}. Find the mean, median, Q1Q_{1}, Q3Q_{3}, range and IQR⁡\operatorname{IQR}. Use our median-of-halves method, excluding the overall median for odd nn.

Need a hint?

The total is 140 min140 \,\mathrm{min}. The 4th4\text{th} value is the median. Split the remaining values into 22 groups of 33.

Show the solution

Mean=1407=20 min\text{Mean}=\frac{140}{7}=20\,\mathrm{min}; median=20 min\text{median} = 20 \,\mathrm{min}. Lower half 1212, 1515, 1818 gives Q1=15 minQ_{1} = 15 \,\mathrm{min}; upper half 2020, 2424, 3131 gives Q3=24 minQ_{3} = 24 \,\mathrm{min}. Range=31−12=19 min\text{Range}=31-12=19\,\mathrm{min}; IQR⁡=24−15=9 min\operatorname{IQR}=24-15=9\,\mathrm{min}.

The mean and middle recorded reading time are both 20 min20 \,\mathrm{min}, and the middle half spans an interval 9 min9 \,\mathrm{min} wide.

2. Sample or population output?

A sample of 3 tasks3\text{ tasks} takes 33, 55 and 7 min7 \,\mathrm{min}. Find s2s^{2} and ss. A calculator shows Sx=2S_x = 2 and σx≈1.633\sigma_x \approx 1.633. Which output answers the question?

Need a hint?

Find the mean, square deviations of −2-2, 00 and 22, and divide their total by 3−13-1.

Show the solution

xˉ=5 min\bar{x} = 5 \,\mathrm{min}. The squared deviations total 4+0+4=8 min24+0+4=8\,\mathrm{min}^{2}. Sample variance=82=4 min2\text{Sample variance}=\frac{8}{2}=4\,\mathrm{min}^{2}; sample standard deviation=4=2 min\text{sample standard deviation}=\sqrt{4}=2\,\mathrm{min}. Use SxS_x. The σx\sigma_x output divides by 33 and corresponds to the population calculation.

3. Percentile wording and ties

1010 students score 5050, 5555, 6060, 6565, 7070, 7070, 7575, 8080, 8585 and 9090 points. What percentage scored at or below 7070? What percentage scored strictly below 7070? Explain why those answers differ.

Need a hint?

Include both 7070-point scores for “at or below”; exclude both for “strictly below.”

Show the solution

At or below 7070: 610×100%=60%\frac{6}{10}\times100\%=60\%. Strictly below 7070: 410×100%=40%\frac{4}{10}\times100\%=40\%. The 22 tied scores account for the difference. Using the inclusive empirical convention, 7070 has percentile rank 6060 within this group. That describes standing among students, rather than the percentage of possible test marks earned.

4. Exactly on an outlier fence

A temperature dataset has Q1=10 ∘CQ_{1} = 10\,{}^{\circ}\mathrm{C} and Q3=20 ∘CQ_{3} = 20\,{}^{\circ}\mathrm{C}. Among readings of −5 ∘C-5\,{}^{\circ}\mathrm{C}, 35 ∘C35\,{}^{\circ}\mathrm{C} and 36 ∘C36\,{}^{\circ}\mathrm{C}, which are flagged by the 1.5×IQR⁡1.5\times\operatorname{IQR} rule?

Need a hint?

IQR⁡=10 ∘C\operatorname{IQR} = 10\,{}^{\circ}\mathrm{C}. Find both fences and apply strict inequalities.

Show the solution

Lower fence=10−1.5(10)=−5 ∘C\text{Lower fence}=10-1.5(10)=-5\,{}^{\circ}\mathrm{C}. Upper fence=20+1.5(10)=35 ∘C\text{Upper fence}=20+1.5(10)=35\,{}^{\circ}\mathrm{C}. Only 36 ∘C36\,{}^{\circ}\mathrm{C} is beyond a fence. −5 ∘C-5\,{}^{\circ}\mathrm{C} and 35 ∘C35\,{}^{\circ}\mathrm{C} are exactly on the fences and are not flagged by this rule.

5. Explain a disagreement between rules

For the Cedar High travel times, Q1=14.5Q_{1} = 14.5, Q3=27Q_{3} = 27, xˉ=21.35\bar{x} = 21.35 and s≈8.8691s \approx 8.8691, all in minutes. Check the 4242-minute journey using each rule. Does a flag prove a recording error?

Need a hint?

Compare 4242 with Q3+1.5IQR⁡Q_3+1.5\operatorname{IQR} and xˉ+2s\bar{x}+2s.

Show the solution

IQR⁡=12.5 min\operatorname{IQR} = 12.5 \,\mathrm{min}; upper IQR fence=45.75 min\text{upper IQR fence}=45.75\,\mathrm{min}. Since 4242 is below that fence, it is not flagged by this rule. The upper 2s2s boundary is about 39.09 min39.09 \,\mathrm{min}, so it is flagged by the more-than-2s2s rule. State the rule with the result. The flag does not prove an error or justify automatic deletion.

6. Convert temperature summaries

A sample has mean 20 ∘C20\,{}^{\circ}\mathrm{C}, median 18 ∘C18\,{}^{\circ}\mathrm{C}, s=3 ∘Cs = 3\,{}^{\circ}\mathrm{C}, IQR⁡=4 ∘C\operatorname{IQR} = 4\,{}^{\circ}\mathrm{C} and s2=9 (∘C)2s^{2} = 9\,({}^{\circ}\mathrm{C})^{2}. Convert these summaries to Fahrenheit using F=1.8C+32F=1.8C+32.

Need a hint?

Apply the full transformation to centers. Apply only the multiplier to spread, and square the multiplier for variance.

Show the solution

Mean=1.8(20)+32=68 ∘F\text{Mean}=1.8(20)+32=68\,{}^{\circ}\mathrm{F}; median=1.8(18)+32=64.4 ∘F\text{median}=1.8(18)+32=64.4\,{}^{\circ}\mathrm{F}. Standard deviation=1.8(3)=5.4 ∘F\text{Standard deviation}=1.8(3)=5.4\,{}^{\circ}\mathrm{F}; IQR⁡=1.8(4)=7.2 ∘F\operatorname{IQR}=1.8(4)=7.2\,{}^{\circ}\mathrm{F}; variance=1.82(9)=29.16 (∘F)2\text{variance}=1.8^{2}(9)=29.16\,({}^{\circ}\mathrm{F})^{2}. Adding 3232 changes location but not spread.

7. Compare without overclaiming

Office A’s sampled processing times have median 10 min10 \,\mathrm{min} and IQR⁡\operatorname{IQR} 4 min4 \,\mathrm{min}. Office B’s have median 13 min13 \,\mathrm{min} and IQR⁡\operatorname{IQR} 8 min8 \,\mathrm{min}. Compare their center and spread. Can you conclude every B application took longer, or that all B applications take longer?

Need a hint?

Use “higher” and “greater variability” with numerical evidence. Separate sample summaries from claims about every observation or a population.

Show the solution

B’s sampled processing times have a 33-minute higher median and an IQR⁡\operatorname{IQR} twice as large as A’s. The middle half of B’s times therefore covers a wider interval. These summaries do not imply every B observation is larger. Generalizing to all applications also needs information about how the samples were obtained.

8. Choose summaries and justify

For the service-time sensitivity example, changing the largest value from 1111 to 25 min25 \,\mathrm{min} moves the mean from 7.337.33 to 8.898.89 and ss from 2.352.35 to 6.33 min6.33 \,\mathrm{min}. The median stays 77 and IQR⁡\operatorname{IQR} stays 4 min4 \,\mathrm{min}. Which center and spread would you report for the alternative distribution, and why?

Need a hint?

Look at the isolated high value and connect it to resistance.

Show the solution

I would report the median of 7 min7 \,\mathrm{min} and IQR⁡\operatorname{IQR} of 4 min4 \,\mathrm{min} because the alternative service-time distribution has an isolated high value, and these summaries are resistant to its effect. The changes in the mean and standard deviation illustrate their sensitivity. This example does not mean resistant summaries can never change.

Quick revision checklist

Choose the tool that answers the question.
TaskToolCheck before reporting
Find an averagexˉ=∑i=1nxin\bar{x}=\frac{\sum_{i=1}^{n}x_i}{n}Total, sample size and units.
Find the middle and quartilesSort; locate middle positions and half medians.Odd/even nn, ties and stated quartile convention.
Describe relative standingPercentile or percentage at/below a cutoff.Reference group and ≤ versus < wording\le\text{ versus }<\text{ wording}.
Find spreadRange; IQR⁡\operatorname{IQR}; s2=∑i=1n(xi−xˉ)2n−1s^{2}=\frac{\sum_{i=1}^{n}(x_i-\bar{x})^{2}}{n-1}; s=s2s=\sqrt{s^{2}}.Spread units; variance units squared; sample output.
Check possible outliersQ1−1.5IQR⁡,Q3+1.5IQR⁡Q_1-1.5\operatorname{IQR},\quad Q_3+1.5\operatorname{IQR}; or xˉ±2s\bar{x}\pm2s.Specified rule, strict inequality and rounding.
Change units: y=a+bxy=a+bx, b>0b>0Centers: a+b×summarya+b\times\text{summary}; spreads: b×summaryb\times\text{summary}; variance: b2×varianceb^{2}\times\text{variance}.Do not add a to a spread\text{Do not add }a\text{ to a spread}.
Choose a center/spread pairMean with ss for reasonably symmetric data without strong outliers; median with IQR⁡\operatorname{IQR} for strong skew or outliers.Support the choice with a distribution feature.
Compare independent samplesDirectly compare the same statistic, with differences or ratios.Variable, group, units and limits of the summaries.

Before submitting: Have I identified the statistic, used the right units, shown enough calculation, answered in context and avoided a stronger claim than the data support?

Final understanding check

88 sampled visitors have recorded waits of 22, 44, 44, 66, 88, 88, 1010 and 22 min22 \,\mathrm{min}.

88 visitor waits: calculate, then interpret

11 dot represents 11 sampled visitor · n=8n = 8

Final-check visitor wait timesEight waits: 2,4,4,6,8,8,10,22 minutes. The largest wait is isolated from the other seven. 0 4 8 12 16 20 24 Recorded visitor wait (minutes)

On a small screen, scroll the plot sideways to read the full scale.

Use the listed values for exact calculations. The isolated 2222-minute observation suggests checking an outlier rule. Invented teaching data.
  1. Calculate the mean, median, quartiles, range and IQR⁡\operatorname{IQR} using our stated method.
  2. The squared deviations from the mean total 272 min2272 \,\mathrm{min}^{2}. Find s2s^{2} and ss.
  3. Apply the 1.5×IQR⁡1.5\times\operatorname{IQR} rule. Choose a center/spread pair and justify your choice.
  4. What percentage of waits is at or below 8 min8 \,\mathrm{min}? Convert the median and IQR⁡\operatorname{IQR} to seconds.
Show the full check and explanation

1. The total is 64 min64 \,\mathrm{min}, so mean=648=8 min\text{mean}=\frac{64}{8}=8\,\mathrm{min}. Median=6+82=7 min\text{Median}=\frac{6+8}{2}=7\,\mathrm{min}. Lower half 22, 44, 44, 66 gives Q1=4 minQ_{1} = 4 \,\mathrm{min}; upper half 88, 88, 1010, 2222 gives Q3=9 minQ_{3} = 9 \,\mathrm{min}. Range=22−2=20 min\text{Range}=22-2=20\,\mathrm{min}; IQR⁡=9−4=5 min\operatorname{IQR}=9-4=5\,\mathrm{min}.

2. s2=2728−1≈38.86 min2s^{2}=\frac{272}{8-1}\approx38.86\,\mathrm{min}^{2}; s=2727≈6.23 mins=\sqrt{\frac{272}{7}}\approx6.23\,\mathrm{min}.

3. Lower IQR fence=4−1.5(5)=−3.5\text{Lower IQR fence}=4-1.5(5)=-3.5 and 9+1.5(5)=16.5 min9+1.5(5)=16.5\,\mathrm{min}. The 2222-minute wait exceeds the upper fence and is flagged. The median of 7 min7 \,\mathrm{min} and IQR⁡\operatorname{IQR} of 5 min5 \,\mathrm{min} are useful because the distribution has an isolated high value and these summaries are resistant.

4. 6 of 8 waits6\text{ of }8\text{ waits} are at or below 8 min8 \,\mathrm{min}: 75%75\%. Multiplying by 6060 gives a median of 420 s420 \,\mathrm{s} and IQR⁡\operatorname{IQR} of 300 s300 \,\mathrm{s}.

Model contextual summary: For these 88 sampled visitors, the middle recorded wait is 7 min7 \,\mathrm{min} and the middle half spans about 4–9 min4\text{–}9\,\mathrm{min}. The 2222-minute wait is flagged by the 1.5×IQR⁡1.5\times\operatorname{IQR} rule. These statements describe this recorded sample; they do not identify the reason for the long wait.

Ready to move on? You should be able to calculate a summary, interpret it with units, apply a stated outlier rule, transform units and justify a comparison or a choice of statistics.

Continue learning

Review quartiles and percentiles · Review spread formulas · Back to the lesson overview