Simple Linear Regression

Fitting a straight line to predict one variable from another by minimizing the sum of squared residuals.

Prerequisites: Correlation, Variance and Standard Deviation.

Simple linear regression fits a straight line through a scatter plot so that one variable, yy, can be described or predicted from another, xx. The line is chosen by least squares: it makes the vertical distances from the points to the line, squared and added up, as small as possible.

It is the simplest statistical model of how one quantity changes with another, and the starting point for nearly all of regression modelling. The fitted slope answers a concrete question, such as “how many extra exam points go with one extra hour of study, on average?”, and the same ideas (residuals, least squares, R2R^2) carry over to models with many predictors.

Intuition

Plot exam scores yy against hours studied xx for a class, and the points rise roughly along a line, with scatter around it. Many lines could be drawn through the cloud. To pick one, we measure how badly each candidate line misses: for each student, the residual is the vertical gap between the observed score and the score the line predicts. Least squares picks the line whose squared residuals have the smallest total.

Eight points rising from about 46 to 78 as hours studied go from 1 to 8, with the fitted line y-hat = 47 + 4x running through them. Short dotted vertical segments join each point to the line; some points lie above the line and some below.
Exam score against hours studied for the eight students in the worked example, with the least-squares line ŷ = 47 + 4x. Each dotted segment is a residual: observed score minus fitted score. The residuals are −5, 4, 2, −1, 4, −4, 1, −1, and their squares sum to 80, the smallest possible for any straight line.

Why vertical distances? Because the goal is to predict yy from xx: the residual is the prediction error in the units of yy. Why squares? Squaring makes every error count as positive, penalizes large misses heavily, and, as shown below, leads to simple closed-form answers.

The model

We observe nn pairs (x1,y1),…,(xn,yn)(x_1, y_1), \dots, (x_n, y_n). Simple linear regression assumes

yi=β0+β1xi+εi,i=1,…,n,y_i = \beta_0 + \beta_1 x_i + \varepsilon_i, \qquad i = 1, \dots, n,

where:

The parameters β0\beta_0 and β1\beta_1 are unknown population quantities. We estimate them from the data by β^0\hat{\beta}_0 and β^1\hat{\beta}_1 (the hats mark estimates), giving the fitted line

y^=β^0+β^1x.\hat{y} = \hat{\beta}_0 + \hat{\beta}_1 x.

Assumptions

The usual assumptions, roughly in order of importance:

  1. Linearity. The mean of yy really does change linearly with xx: E⁡[εi]=0\E[\varepsilon_i] = 0 for every xix_i. If the true relationship is curved, the line is a poor summary no matter how it is fitted.
  2. Independent errors. The errors for different observations are independent. This fails, for example, for repeated measurements on the same person or for data collected over time.
  3. Constant variance. The errors have the same variance σ2\sigma^2 at every value of xx. If the scatter fans out as xx grows, this fails.
  4. Normal errors. The errors follow a normal distribution, εi∼N(0,σ2)\varepsilon_i \sim \mathcal{N}(0, \sigma^2).

The least-squares line itself can be computed without any of these assumptions; they matter for interpreting it and for inference. Normality is needed only for exact confidence intervals and hypothesis tests in small samples. With larger samples, those procedures remain approximately valid without it, by the central limit theorem. The other three assumptions cannot be rescued by a larger sample.

Least squares

For a candidate intercept b0b_0 and slope b1b_1, the residual of point ii is yi−(b0+b1xi)y_i - (b_0 + b_1 x_i). The sum of squared residuals is

SSE(b0,b1)=∑i=1n(yi−b0−b1xi)2.\text{SSE}(b_0, b_1) = \sum_{i=1}^n \big(y_i - b_0 - b_1 x_i\big)^2.

The least-squares estimates β^0,β^1\hat{\beta}_0, \hat{\beta}_1 are the values of b0,b1b_0, b_1 that minimize it. Writing

Sxx=∑i=1n(xi−xˉ)2,Sxy=∑i=1n(xi−xˉ)(yi−yˉ),S_{xx} = \sum_{i=1}^n (x_i - \bar{x})^2, \qquad S_{xy} = \sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y}),

the solution is

β^1=SxySxx,β^0=yˉ−β^1xˉ.\hat{\beta}_1 = \frac{S_{xy}}{S_{xx}}, \qquad \hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x}.

This requires Sxx>0S_{xx} > 0, that is, the xx values are not all equal.

Where the formulas come from

SSE is a smooth bowl-shaped function of b0b_0 and b1b_1, so its minimum is where both partial derivatives are zero. Writing ei=yi−b0−b1xie_i = y_i - b_0 - b_1 x_i, the two conditions are

∑i=1nei=0,∑i=1nxiei=0.\sum_{i=1}^n e_i = 0, \qquad \sum_{i=1}^n x_i e_i = 0.

The first condition says the residuals sum to zero. Dividing it by nn gives yˉ−b0−b1xˉ=0\bar{y} - b_0 - b_1\bar{x} = 0, so b0=yˉ−b1xˉb_0 = \bar{y} - b_1 \bar{x}: the least-squares line always passes through the point of means (xˉ,yˉ)(\bar{x}, \bar{y}).

Substituting this into eie_i gives ei=(yi−yˉ)−b1(xi−xˉ)e_i = (y_i - \bar{y}) - b_1(x_i - \bar{x}). Because ∑ei=0\sum e_i = 0, the second condition can be rewritten as ∑(xi−xˉ)ei=0\sum (x_i - \bar{x}) e_i = 0, which becomes

∑i=1n(xi−xˉ)(yi−yˉ)−b1∑i=1n(xi−xˉ)2=0,\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y}) - b_1 \sum_{i=1}^n (x_i - \bar{x})^2 = 0,

that is, Sxy−b1Sxx=0S_{xy} - b_1 S_{xx} = 0, so b1=Sxy/Sxxb_1 = S_{xy}/S_{xx}.

The slope in terms of correlation

Dividing the top and bottom of Sxy/SxxS_{xy}/S_{xx} by n−1n - 1 turns them into the sample covariance sxys_{xy} and the sample variance sx2s_x^2. Since the sample correlation is r=sxy/(sxsy)r = s_{xy}/(s_x s_y),

β^1=sxysx2=r sysx.\hat{\beta}_1 = \frac{s_{xy}}{s_x^2} = r\,\frac{s_y}{s_x}.

The slope is the correlation, rescaled from “standard deviations of yy per standard deviation of xx” into the actual units of yy per unit of xx.

Interpreting the fit

R² and the fraction of variance explained

The total variation of yy around its mean, SST=∑(yi−yˉ)2\text{SST} = \sum (y_i - \bar{y})^2, splits into a part explained by the line and a part left in the residuals:

∑i=1n(yi−yˉ)2⏟SST=∑i=1n(y^i−yˉ)2⏟explained+∑i=1n(yi−y^i)2⏟SSE.\underbrace{\sum_{i=1}^n (y_i - \bar{y})^2}_{\text{SST}} = \underbrace{\sum_{i=1}^n (\hat{y}_i - \bar{y})^2}_{\text{explained}} + \underbrace{\sum_{i=1}^n (y_i - \hat{y}_i)^2}_{\text{SSE}}.

The coefficient of determination is the explained fraction:

R2=1−SSESST.R^2 = 1 - \frac{\text{SSE}}{\text{SST}}.

It lies between 00 (the line does no better than predicting yˉ\bar{y} for everyone) and 11 (every point lies on the line). In simple linear regression, R2=r2R^2 = r^2, the square of the sample correlation between xx and yy.

Worked example

Eight students record how many hours xx they studied for an exam and their score yy out of 100.

xix_i yiy_i xi−xˉx_i - \bar{x} yi−yˉy_i - \bar{y} (xi−xˉ)2(x_i - \bar{x})^2 (xi−xˉ)(yi−yˉ)(x_i - \bar{x})(y_i - \bar{y})
1 46 −3.5-3.5 −19-19 12.25 66.5
2 59 −2.5-2.5 −6-6 6.25 15
3 61 −1.5-1.5 −4-4 2.25 6
4 62 −0.5-0.5 −3-3 0.25 1.5
5 71 0.50.5 66 0.25 3
6 67 1.51.5 22 2.25 3
7 76 2.52.5 1111 6.25 27.5
8 78 3.53.5 1313 12.25 45.5
Sum 42 168

Step 1: means. xˉ=36/8=4.5\bar{x} = 36/8 = 4.5 hours and yˉ=520/8=65\bar{y} = 520/8 = 65 points.

Step 2: slope. From the column sums, Sxx=42S_{xx} = 42 and Sxy=168S_{xy} = 168, so

β^1=16842=4 points per hour.\hat{\beta}_1 = \frac{168}{42} = 4 \text{ points per hour}.

Step 3: intercept.

β^0=yˉ−β^1xˉ=65−4×4.5=47 points.\hat{\beta}_0 = \bar{y} - \hat{\beta}_1 \bar{x} = 65 - 4 \times 4.5 = 47 \text{ points}.

The fitted line is y^=47+4x\hat{y} = 47 + 4x. Each additional hour of study is associated with 4 more points on average. The intercept, 47, is the predicted score for a student who did not study at all; that is slightly outside the data (the minimum was 1 hour), so treat it with some caution.

Step 4: residuals. The fitted values at x=1,…,8x = 1, \dots, 8 are 51,55,59,63,67,71,75,7951, 55, 59, 63, 67, 71, 75, 79, so the residuals are

−5, 4, 2, −1, 4, −4, 1, −1.-5,\ 4,\ 2,\ -1,\ 4,\ -4,\ 1,\ -1.

They sum to zero, as least squares guarantees. Their squares sum to SSE=80\text{SSE} = 80, and the residual standard deviation is se=80/6≈3.65s_e = \sqrt{80/6} \approx 3.65 points: a typical student’s score lies within a few points of the line.

Step 5: R². The squared deviations (yi−yˉ)2(y_i - \bar{y})^2 sum to SST=752\text{SST} = 752. So

R2=1−80752≈0.894.R^2 = 1 - \frac{80}{752} \approx 0.894.

About 89% of the variation in scores among these students is accounted for by the linear relationship with hours studied. As a check, r=168/42×752≈0.945r = 168/\sqrt{42 \times 752} \approx 0.945, and r2≈0.894r^2 \approx 0.894. Also sx=42/7≈2.45s_x = \sqrt{42/7} \approx 2.45 and sy=752/7≈10.36s_y = \sqrt{752/7} \approx 10.36, and r sy/sx≈0.945×10.36/2.45≈4r\, s_y/s_x \approx 0.945 \times 10.36 / 2.45 \approx 4, matching the slope.

Step 6: predict. For a student who studies 5.5 hours, the line predicts 47+4×5.5=6947 + 4 \times 5.5 = 69 points.

Checking the assumptions with residual plots

The scatter plot with the fitted line is the first check. The second is a residual plot: residuals eie_i on the vertical axis against xix_i (or against the fitted values y^i\hat{y}_i). If the assumptions hold, the residuals form a shapeless horizontal band around zero. Typical warning signs:

A histogram or normal quantile plot of the residuals can check normality, but with small samples such checks have little power, and with large samples normality matters less.

Regression to the mean

Dividing both sides of the fitted line by the standard deviations shows that, in z-score units, the predicted yy is

y^−yˉsy=r x−xˉsx.\frac{\hat{y} - \bar{y}}{s_y} = r \, \frac{x - \bar{x}}{s_x}.

Unless the points lie exactly on a line, ∣r∣<1|r| < 1, so the prediction is always fewer standard deviations from the mean than xx is. In the example, a student who studies one standard deviation above average (about 6.95 hours) is predicted to score 0.9450.945 standard deviations above average, not a full one.

This is regression to the mean, and it is where the word “regression” comes from: Francis Galton observed in the 19th century that unusually tall parents tend to have children who are tall, but less extremely so. It is a statistical effect, not a causal force. It also explains many apparent “treatment effects”: a group selected because it scored unusually badly will, on average, score closer to the mean next time even if nothing is done.

Limitations

Extrapolation. The line describes the data in the range where it was observed. Outside that range there is no evidence that the relationship stays linear. Our line predicts 47+4×20=12747 + 4 \times 20 = 127 points for a student who studies 20 hours, which is impossible on a test out of 100. Predictions far from the observed xx values deserve little trust.

Causation. A regression of yy on xx describes association, just like correlation. In the example, students who study more may also attend more classes or have more prior knowledge, and these confounders could explain part of the slope. Interpreting β^1\hat{\beta}_1 as “the effect of one more hour of study” needs a randomized experiment or a careful argument that confounding is absent.

Direction matters. Regressing yy on xx and regressing xx on yy give different lines (unless r=±1r = \pm 1), because one minimizes vertical distances and the other horizontal ones. Choose the variable you want to predict as yy.

Common misunderstandings

“A high R² means the model is correct.” A curved relationship can still give a high R2R^2 for a straight line, and a line with high R2R^2 can still violate the assumptions. Look at the residual plot.

“A low R² means x has no effect.” A small R2R^2 means xx explains little of the variation in yy, but the slope can still be real and important, especially when yy has many other influences.

“The slope tells us what happens when we change x.” Only under causal assumptions. Without them it describes how the average of yy differs between groups with different xx.

“Normality of the data is required to fit a regression.” Least squares can be computed for any data. Normality concerns the errors, not xx or yy themselves, and it matters only for exact small-sample inference.

Further reading