Line of Best Fit
Through any scatter of points there is exactly one line that misses them by the least, measured the least squares way. Drag the points and it moves to stay the best.
Last updated
Plot any real data and the points refuse to sit on one line. You can still draw a line through the cloud that captures the trend, and different people with rulers will draw slightly different ones. The line of best fit settles the argument: there is one line that is best by a precise measure, and a formula that finds it.
Drag any point below and the line moves to stay the best fit. Pull a point far from the others and notice how hard it pulls the line towards itself.
Drag any point and the line moves to stay the best fit. Points far from the rest pull it the hardest.
What "best" means
A line misses each point by a vertical distance: the point's actual y minus the y the line predicts at that x. That distance is the point's residual. A good line should make the residuals small, but you cannot just add them up, because misses above the line are positive and misses below are negative and they cancel.
So each residual is squared first, which makes every miss count and makes a big miss count a lot more than a small one. The line of best fit is the line that makes the sum of the squared residuals as small as it can be. That is why it is called the least squares line, and for any set of points whose x values are not all the same there is exactly one of it.
The two formulas
Write the line as y = mx + b, with the means of the data written x̄ and ȳ. Minimising the sum of squares gives the slope and intercept directly:
m = Σ(x - x̄)(y - ȳ) / Σ(x - x̄)²
b = ȳ - m x̄
The top of the slope formula is the same sum of products that sits on top of the correlation coefficient, which is why the slope and r always have the same sign. The bottom only involves x: it measures how spread out the x values are.
A worked example
Five points: x = 2, 4, 6, 8, 10 and y = 3, 7, 8, 12, 15. The mean of x is 6 and the mean of y is 9.
| x | y | x - 6 | y - 9 | product | (x - 6)² |
|---|---|---|---|---|---|
| 2 | 3 | -4 | -6 | 24 | 16 |
| 4 | 7 | -2 | -2 | 4 | 4 |
| 6 | 8 | 0 | -1 | 0 | 0 |
| 8 | 12 | 2 | 3 | 6 | 4 |
| 10 | 15 | 4 | 6 | 24 | 16 |
| total | 58 | 40 |
The slope is m = 58/40 = 29/20, which is 1.45. For the intercept, put the point of means into the line and solve for b:
So b = 3/10 and the line of best fit is y = 1.45x + 0.3. The linear regression calculator does the same work with the running-totals form of the formulas and shows each sum.
It always passes through the point of means
Rearrange the intercept formula and you get ȳ = m x̄ + b. That says the point (x̄, ȳ) sits exactly on the line, whatever the data. In the example, 1.45 × 6 + 0.3 = 9.
This is a useful check on any line you calculate, and a useful start when you draw one by eye: mark the point of means first, then pivot the ruler around it until the points are balanced on either side.
Passing through the point of means is not enough on its own, though. For x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5 the point of means is (3, 4) and the least squares line is y = 0.6x + 2.2, with a sum of squared residuals of 2.4. The lines y = 0.5x + 2.5 and y = x + 1 both pass through (3, 4) as well, and their sums of squares are 2.5 and 4. Among all lines through that point, only slope 0.6 gets down to 2.4.
Predicting with the line
Once you have the line, a prediction is a substitution. With y = 1.45x + 0.3, an x of 7 gives 1.45 × 7 + 0.3 = 10.45.
That prediction is an interpolation: 7 lies inside the data, between the points at 6 and 8, so there is evidence on both sides that the trend holds there. A prediction outside the range of the data is an extrapolation, and nothing in the data supports it.
A concrete case. Six students studied for 1 to 6 hours and scored 52, 58, 61, 67, 70 and 76. The line of best fit is y = 162/35x + 47.8, a slope of about 4.63 points per hour, with r ≈ 0.9959. Inside the data it predicts well: 4.5 hours gives about 68.63. Extrapolate to 12 hours and it predicts about 103.34, on a test marked out of 100. The line is fine; the trend it describes simply stops somewhere past the data, and the line cannot know where.
Where a straight line is the wrong shape
A line of best fit always exists, even when a straight line is the wrong model. The squares 0, 1, 4, 9, 16, 25 at x = 0 to 5 have a line of best fit y = 5x - 10/3 with r ≈ 0.9599, which sounds excellent, and yet the points lie on a curve. The residuals give it away: positive at both ends and negative in the middle. The residuals page shows how to read that pattern, and the quadratic regression calculator fits the curve instead.
Try one
The actual point at x = 4 is (4, 7), so the line misses it by 0.9. That miss is the point's residual, and the sum of all five squared misses, 1.9, is the smallest any straight line can achieve for these points.
Common questions
- What is a line of best fit?
- A straight line drawn through a scatter plot to summarise the trend in the points. The line of best fit in the precise sense is the least squares regression line: of all possible straight lines, the one whose vertical misses from the points, squared and added up, give the smallest total.
- How do you find the line of best fit?
- Find the mean of x and the mean of y. The slope is the sum of (x minus mean of x) times (y minus mean of y), divided by the sum of (x minus mean of x) squared. The intercept is the mean of y minus the slope times the mean of x. For x = 2, 4, 6, 8, 10 and y = 3, 7, 8, 12, 15 that gives a slope of 58/40 = 1.45 and an intercept of 9 minus 1.45 × 6 = 0.3, so y = 1.45x + 0.3.
- Does the line of best fit have to go through the origin?
- No. It passes through the point of means, not the origin. Forcing it through (0, 0) is a different model, used only when there is a reason y must be 0 when x is 0, and it generally fits the points worse.
- Does the line of best fit have to pass through any of the points?
- No, and usually it passes through none of them. The one point it always passes through is (mean of x, mean of y), and that point is often not one of the data points. A line drawn to touch as many points as possible is not the least squares line.
- What is the difference between interpolation and extrapolation?
- Interpolation is predicting y for an x inside the range of the data, where the line has points on both sides to support it. Extrapolation is predicting outside that range, where nothing in the data says the trend continues. Interpolation is generally trustworthy when the fit is good; extrapolation can be badly wrong even with r close to 1.
- Is a line of best fit the same as linear regression?
- Yes, in the usual sense. Linear regression is the method, and the least squares line it produces is the line of best fit. A line drawn by eye on a graph is also called a line of best fit in school, but it is an estimate of the regression line, not the line itself.
Want to run this on your own numbers?
Linear Regression Calculator Open the calculator