Simple Linear Regression
Fitting a straight line to predict one variable from another by minimizing the sum of squared residuals.
Prerequisites: Correlation, Variance and Standard Deviation.
Simple linear regression fits a straight line through a scatter plot so that one variable, , can be described or predicted from another, . The line is chosen by least squares: it makes the vertical distances from the points to the line, squared and added up, as small as possible.
It is the simplest statistical model of how one quantity changes with another, and the starting point for nearly all of regression modelling. The fitted slope answers a concrete question, such as “how many extra exam points go with one extra hour of study, on average?”, and the same ideas (residuals, least squares, ) carry over to models with many predictors.
Intuition
Plot exam scores against hours studied for a class, and the points rise roughly along a line, with scatter around it. Many lines could be drawn through the cloud. To pick one, we measure how badly each candidate line misses: for each student, the residual is the vertical gap between the observed score and the score the line predicts. Least squares picks the line whose squared residuals have the smallest total.
Why vertical distances? Because the goal is to predict from : the residual is the prediction error in the units of . Why squares? Squaring makes every error count as positive, penalizes large misses heavily, and, as shown below, leads to simple closed-form answers.
The model
We observe pairs . Simple linear regression assumes
where:
- is the intercept: the mean of when ;
- is the slope: the change in the mean of for a one-unit increase in ;
- is the error for observation : everything about that the line does not capture, such as other influences and measurement noise.
The parameters and are unknown population quantities. We estimate them from the data by and (the hats mark estimates), giving the fitted line
Assumptions
The usual assumptions, roughly in order of importance:
- Linearity. The mean of really does change linearly with : for every . If the true relationship is curved, the line is a poor summary no matter how it is fitted.
- Independent errors. The errors for different observations are independent. This fails, for example, for repeated measurements on the same person or for data collected over time.
- Constant variance. The errors have the same variance at every value of . If the scatter fans out as grows, this fails.
- Normal errors. The errors follow a normal distribution, .
The least-squares line itself can be computed without any of these assumptions; they matter for interpreting it and for inference. Normality is needed only for exact confidence intervals and hypothesis tests in small samples. With larger samples, those procedures remain approximately valid without it, by the central limit theorem. The other three assumptions cannot be rescued by a larger sample.
Least squares
For a candidate intercept and slope , the residual of point is . The sum of squared residuals is
The least-squares estimates are the values of that minimize it. Writing
the solution is
This requires , that is, the values are not all equal.
Where the formulas come from
SSE is a smooth bowl-shaped function of and , so its minimum is where both partial derivatives are zero. Writing , the two conditions are
The first condition says the residuals sum to zero. Dividing it by gives , so : the least-squares line always passes through the point of means .
Substituting this into gives . Because , the second condition can be rewritten as , which becomes
that is, , so .
The slope in terms of correlation
Dividing the top and bottom of by turns them into the sample covariance and the sample variance . Since the sample correlation is ,
The slope is the correlation, rescaled from “standard deviations of per standard deviation of ” into the actual units of per unit of .
Interpreting the fit
- Slope. is the estimated change in the average of associated with a one-unit increase in . It describes the trend across the population, not what will happen to any individual.
- Intercept. is the predicted at . It is meaningful only if is a sensible value inside or near the range of the data; otherwise it is just the number that positions the line.
- Residuals. The residuals are what the line leaves unexplained. Their typical size is summarized by the residual standard deviation , which estimates the error standard deviation . The divisor is because two parameters were estimated from the data.
R² and the fraction of variance explained
The total variation of around its mean, , splits into a part explained by the line and a part left in the residuals:
The coefficient of determination is the explained fraction:
It lies between (the line does no better than predicting for everyone) and (every point lies on the line). In simple linear regression, , the square of the sample correlation between and .
Worked example
Eight students record how many hours they studied for an exam and their score out of 100.
| 1 | 46 | 12.25 | 66.5 | ||
| 2 | 59 | 6.25 | 15 | ||
| 3 | 61 | 2.25 | 6 | ||
| 4 | 62 | 0.25 | 1.5 | ||
| 5 | 71 | 0.25 | 3 | ||
| 6 | 67 | 2.25 | 3 | ||
| 7 | 76 | 6.25 | 27.5 | ||
| 8 | 78 | 12.25 | 45.5 | ||
| Sum | 42 | 168 |
Step 1: means. hours and points.
Step 2: slope. From the column sums, and , so
Step 3: intercept.
The fitted line is . Each additional hour of study is associated with 4 more points on average. The intercept, 47, is the predicted score for a student who did not study at all; that is slightly outside the data (the minimum was 1 hour), so treat it with some caution.
Step 4: residuals. The fitted values at are , so the residuals are
They sum to zero, as least squares guarantees. Their squares sum to , and the residual standard deviation is points: a typical student’s score lies within a few points of the line.
Step 5: R². The squared deviations sum to . So
About 89% of the variation in scores among these students is accounted for by the linear relationship with hours studied. As a check, , and . Also and , and , matching the slope.
Step 6: predict. For a student who studies 5.5 hours, the line predicts points.
Checking the assumptions with residual plots
The scatter plot with the fitted line is the first check. The second is a residual plot: residuals on the vertical axis against (or against the fitted values ). If the assumptions hold, the residuals form a shapeless horizontal band around zero. Typical warning signs:
- A curve (residuals positive at both ends and negative in the middle, or the reverse): the relationship is not linear.
- A funnel (residuals spreading out as grows): the variance is not constant.
- One or two isolated points far from zero: outliers, which can pull the line towards themselves, much as one point can distort a correlation.
- Trends in the order the data were collected: the errors may not be independent.
A histogram or normal quantile plot of the residuals can check normality, but with small samples such checks have little power, and with large samples normality matters less.
Regression to the mean
Dividing both sides of the fitted line by the standard deviations shows that, in z-score units, the predicted is
Unless the points lie exactly on a line, , so the prediction is always fewer standard deviations from the mean than is. In the example, a student who studies one standard deviation above average (about 6.95 hours) is predicted to score standard deviations above average, not a full one.
This is regression to the mean, and it is where the word “regression” comes from: Francis Galton observed in the 19th century that unusually tall parents tend to have children who are tall, but less extremely so. It is a statistical effect, not a causal force. It also explains many apparent “treatment effects”: a group selected because it scored unusually badly will, on average, score closer to the mean next time even if nothing is done.
Limitations
Extrapolation. The line describes the data in the range where it was observed. Outside that range there is no evidence that the relationship stays linear. Our line predicts points for a student who studies 20 hours, which is impossible on a test out of 100. Predictions far from the observed values deserve little trust.
Causation. A regression of on describes association, just like correlation. In the example, students who study more may also attend more classes or have more prior knowledge, and these confounders could explain part of the slope. Interpreting as “the effect of one more hour of study” needs a randomized experiment or a careful argument that confounding is absent.
Direction matters. Regressing on and regressing on give different lines (unless ), because one minimizes vertical distances and the other horizontal ones. Choose the variable you want to predict as .
Common misunderstandings
“A high R² means the model is correct.” A curved relationship can still give a high for a straight line, and a line with high can still violate the assumptions. Look at the residual plot.
“A low R² means x has no effect.” A small means explains little of the variation in , but the slope can still be real and important, especially when has many other influences.
“The slope tells us what happens when we change x.” Only under causal assumptions. Without them it describes how the average of differs between groups with different .
“Normality of the data is required to fit a regression.” Least squares can be computed for any data. Normality concerns the errors, not or themselves, and it matters only for exact small-sample inference.
Further reading
- David Freedman, Robert Pisani, and Roger Purves, Statistics, 4th ed., W. W. Norton, 2007. Explains regression, the regression effect, and their misuses with unusual clarity.
- David M. Diez, Mine Çetinkaya-Rundel, and Christopher D. Barr, OpenIntro Statistics, 4th ed., 2019. Free online; covers least squares, residual plots, and inference for the slope.
- NIST/SEMATECH, e-Handbook of Statistical Methods, https://www.itl.nist.gov/div898/handbook/. A practical reference on fitting and checking regression models.