AP Statistics / Unit 5: Regression Analysis / Topic 5.5
NUM8ERS study notes · Topic 5.5

Least-Squares Regression

Many lines can run through a scatterplot. Which one should you use? Learn how least squares chooses a line, use technology to find it, and explain what its coefficients and model fit mean in a real situation.

2026–27 curriculum4 worked examples8 practice questions3 visual explanations

By the end of this lesson, you should be able to:

  • Explain what the least-squares line minimizes and identify the mean point it passes through.
  • Use technology to find the fitted equation and distinguish the intercept, slope and correlation.
  • Interpret slope, intercept and the coefficient of determination in context.

Before you start: Review linear regression models and residuals. A residual is an actual response minus its prediction.

First time learning this? Follow the delivery data from scatterplot to fitted line, then read the output and interpretation examples.

Here to revise? Use the coefficient checklist, then attempt the practice before opening solutions.

The concept in 60 seconds

The least-squares regression line is the line that makes the sum of squared residuals as small as possible for the observed data. Its usual prediction form is:

y^=a+bx.\hat{y}=a+bx.

The intercept aa gives the predicted response when x=0x=0. The slope bb tells you how much the predicted response changes for each additional unit of xx.

We square each vertical prediction error before adding. This prevents positive and negative residuals from canceling and gives larger misses more weight.

Best for a specific purpose: “Best” here means the smallest sum of squared residuals among straight lines. It does not guarantee that a line is an appropriate model or that every prediction is accurate.

Quick check: does the least-squares line have to pass through every point?

No. It balances the squared misses across the whole data set. When an intercept is included, it does pass through the mean point (xˉ,yˉ)(\bar{x},\bar{y}).

A real situation: choosing a delivery-time line

A delivery team records distance xx in kilometres and time yy in minutes for 99 invented deliveries. These are the same teaching data used in Topics 5.3 and 5.4.

Paired observations: keep each distance with its own delivery time.
Delivery IDDistance xx (km)Observed time yy (min)
A552222
B7.57.524.524.5
C10102929
D12.512.535.535.5
E15153939
F17.517.544.544.5
G20204747
H22.522.551.551.5
I25255858

Technology gives the least-squares line:

y^=12+1.8x.\hat{y}=12+1.8x.

The means are xˉ=15\bar{x}=15 km and yˉ=39\bar{y}=39 minutes. The line predicts 12+1.8(15)=3912+1.8(15)=39 minutes at the mean distance, so it passes through (15,39)(15,39).

The fitted line passes through the mean point
image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 5 10 15 20 25 Distance (km) 10 20 30 40 50 60 Time (min) $(\bar{x},\bar{y})=(15,39)$ $\hat{y}=12+1.8x$ Least-squares line Observed deliveries image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 5 10 15 20 25 Distance (km) 10 20 30 40 50 60 Time (min) $(\bar{x},\bar{y})=(15,39)$ $\hat{y}=12+1.8x$ Least-squares line Observed deliveries

Circles are observed deliveries. The line predicts time from distance. The highlighted mean point lies on the fitted line, even though it need not be an observed delivery.

These data were constructed to make the arithmetic easy; their exceptionally close fit is not typical of all real delivery data. Use the observed distance range, 5≤x≤255\leq x\leq25, when judging whether a prediction or interpretation is supported.

Key ideas and notation

Response and prediction

yy is an observed response; y^\hat{y} is the response predicted at the same xx. The hat matters.

Intercept: aa

The predicted response at x=0x=0. Its units are the response units.

Slope: bb

The change in predicted response per unit increase in xx. Its units are response units per explanatory unit.

Residual: ee

The signed vertical difference e=y−y^e=y-\hat{y}. Least squares uses e2e^2 to compare lines.

Correlation: rr

A unitless measure of the direction and strength of linear association. It is not the slope.

Coefficient of determination: r2r^2

The proportion of variation in the response explained by its linear relationship with the explanatory variable.

Watch the output convention: some calculators label slope aa and intercept bb in the form y=ax+by=ax+b. Others use y=a+bxy=a+bx. Identify coefficients from the displayed equation rather than memorizing which letter is the slope.

What least squares minimizes

For each observation, calculate the residual, square it, then add all squared residuals. We call this total the sum of squared errors, or SSE\mathrm{SSE}.

SSE=∑i=1n(yi−y^i)2.\mathrm{SSE}=\sum_{i=1}^{n}(y_i-\hat{y}_i)^2.

The index ii simply labels an observation, and nn is the number of observations. For the delivery line, the residuals in minutes are:

(1,−1,−1,1,0,1,−1,−1,1).(1,-1,-1,1,0,1,-1,-1,1).

There are 88 residuals of magnitude 11 and one of 00, giving SSE=8(12)+02=8\mathrm{SSE}=8(1^2)+0^2=8 square minutes. The signed residuals add to 00, but the squared errors do not.

Compare candidate lines using the same delivery data. Smaller squared-error total is better by the least-squares criterion.
Candidate equationSSE\mathrm{SSE} (min2^2)What changes?
y^=12+1.8x\hat{y}=12+1.8x88Least-squares fit
y^=13+1.8x\hat{y}=13+1.8x1717Same slope, higher intercept
y^=9+2x\hat{y}=9+2x2323Steeper line through the same mean point
Same data, different squared-error totals
image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0 10 20 $\mathrm{SSE}\ (\mathrm{min}^2)$ $\hat{y}=12+1.8x$ $\hat{y}=13+1.8x$ $\hat{y}=9+2x$ $8$ $17$ $23$ image/svg+xml Matplotlib v3.11.2, https://matplotlib.org/ 0 10 20 $\mathrm{SSE}\ (\mathrm{min}^2)$ $8$ $17$ $23$ $\hat{y}=12+1.8x$ $\hat{y}=13+1.8x$ $\hat{y}=9+2x$

The fitted line has the smallest squared-error total. The other candidates illustrate why changing either the height or slope can increase the total. Squared-error units are square minutes.

Checking these candidates illustrates the rule; technology has found the minimum over all straight lines with an intercept. The line y^=9+2x\hat{y}=9+2x also passes through (15,39)(15,39), so passing through the mean point alone does not prove a line is the least-squares line.

Vertical, not perpendicular: errors are measured in the response direction at the same explanatory value. The criterion is not shortest slanted distances, total absolute residuals, or the sum of signed residuals.

With an intercept and nonconstant explanatory values, the fitted line passes through (xˉ,yˉ)(\bar{x},\bar{y}) and its residuals sum to 00, apart from rounding. Swapping xx and yy changes the prediction task and generally produces a different line.

Interpret the slope and intercept

Slope: a change in predicted response

For y^=12+1.8x\hat{y}=12+1.8x, the slope is b=1.8b=1.8 minutes per kilometre.

For each additional kilometre of delivery distance, the model predicts an increase of 1.81.8 minutes in delivery time, on average.

This describes the change in the prediction. It does not mean every pair of deliveries differs by exactly that amount, or that distance is the only cause of travel time. Over an increase of 55 km, the predicted change is 1.8(5)=91.8(5)=9 minutes.

Intercept: the prediction at zero

The intercept is a=12a=12 minutes: the model predicts a delivery time of 1212 minutes at a distance of 00 km.

However, 00 km is outside the observed range 5≤x≤255\leq x\leq25. The intercept helps position the fitted line, but the data do not support interpreting it as an established loading time or a reliable zero-distance delivery time.

Meaningfulness check: Is x=0x=0 sensible in the situation? Is it within or reasonably near the observed range? Is the predicted response logically possible? Explain any limitation instead of inventing a story for the intercept.

What if the slope is negative?

Say the predicted response decreases. For a hypothetical battery model y^=98−6x\hat{y}=98-6x, where xx is hours of use and yy is charge percentage, an additional hour predicts a decrease of 66 percentage points in charge. This is a change in percentage points, not a relative decrease of 6%6\%.

Understand the coefficient of determination

In simple least-squares regression with an intercept, r2r^2 is the square of the correlation. It describes the proportion of variation in observed responses, about their mean, explained by the fitted linear relationship.

Posted on Google Google
0000003998 : Abdad Alam Shamim Alam profile picture
0000003998 : Abdad Alam Shamim Alam
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
This is the bestest institution I have ever come to and I love it very much.
Posted on Google Google
Aliki S profile picture
Aliki S
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Ali profile picture
Ali
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Sly is him
Posted on Google Google
Jitendra Kumar Kumawat profile picture
Jitendra Kumar Kumawat
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Raahil Hasan profile picture
Raahil Hasan
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
Posted on Google Google
Tina Mudarres profile picture
Tina Mudarres
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
The classes are amazing and my child learnt so much
Posted on Google Google
alisha gadoya profile picture
alisha gadoya
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I have experienced a lot of good things, it has taught me so many things and from B/Cs I have gone to an A! This is wonderful and Mavish’s class is awesome.
Posted on Google Google
smasher 123 profile picture
smasher 123
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
It’s sensational my kid went. First he only would get C now it’s all A’s really good and recommend
Posted on Google Google
Mosa Al- Samaraie profile picture
Mosa Al- Samaraie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
AMAZING PRICES GREAT TEACHERS STRAIGHT FORWARD LEARNING STEADY PACE IN TUTORING
Posted on Google Google
Cael Dagnelie profile picture
Cael Dagnelie
Google star 1Google star 2Google star 3Google star 4Google star 5Trustindex verifies that the original source of the review is Google.
I had a great experience learning math with Ms. Mavish. She explains complex topics in a very clear and simple way, which made it easier for me to understand and enjoy the subject. Her patience and dedication really stood out, and she always made sure that everyone in the class was keeping up. I especially appreciated how approachable she was — I never felt afraid to ask questions, and she was always willing to help. Thanks to her teaching, my confidence in math has grown a lot. I’m really thankful for the effort she puts into every lesson!

NUM8ERS is one of finest tutoring institutes in UAE, Located in Al Barsha 1, Dubai. Close to DUBAI AMERICAN ACADEMY (DAA) & AMERICAN SCHOOL OF DUBAI (ASD).