Menu

Linear Regression Calculator

The line of best fit through your points, exact, with r and every step of the working.

By Nethanel Bar, Co-founder & CEO

Last updated

Want to solve these without the calculator?

The Coddy math course teaches the method itself - you work each step on an interactive board and get told exactly where a move went wrong.

Linear regression finds the one line that misses the points by the least

Real data never sits on one straight line, so any line you draw through it misses most of the points. Each miss is a residual: the actual y minus the y the line predicts. Linear regression squares every residual, adds the squares, and picks the one line that makes the total as small as possible. That line is the least squares regression line, also called the line of best fit, and there is exactly one of it for any set of points whose x values are not all the same.

Finding it takes five numbers: the count n and the sums of x, y, xy and x². The slope is m = (nΣxy minus ΣxΣy) divided by (nΣx² minus (Σx)²), and the intercept is b = (Σy minus mΣx) divided by n. This page builds the table of those sums, puts your numbers into both formulas, and shows each line of the working, because the table is the part a teacher checks.

The slope and intercept are kept as exact fractions. On data made of whole numbers, decimals and fractions they are always rational, so a slope of 162/35 is printed as 162/35 and not as 4.6286, which is a rounding of it. The correlation coefficient r is a square root, so it appears as a radical with its decimal beside it. The Copy button gives you the line in decimals, ready for a spreadsheet.

What to notice about the line

  • It always passes through the point of means. For x = 1 to 5 and y = 2, 4, 5, 4, 5 the means are 3 and 4, the line is y = 0.6x + 2.2, and 0.6 × 3 + 2.2 = 4.
  • Its residuals always add up to exactly zero. Misses above the line and below it balance, which is why the method squares them before adding.
  • The slope has units: y units per x unit. A slope of 162/35 on hours studied against test score means about 4.63 more points for each extra hour.
  • Swapping x and y gives a different line, not the same line rearranged. Regressing y on x makes the vertical misses small; regressing x on y makes the horizontal ones small.
  • The slope says how steep the trend is and r says how tightly the points follow it. They always share a sign, but a steep line can have a weak r and a shallow line a strong one.
  • Two points always give a perfect fit with r = 1 or -1. That says nothing about the data; a regression line starts to mean something with three or more points.

How to find the regression line by hand

  1. Make the table

    Write each point as a row with four columns: x, y, xy and x². Add up every column and count the points. For x = 1, 2, 3, 4, 5 and y = 2, 4, 5, 4, 5 that gives n = 5, Σx = 15, Σy = 20, Σxy = 66 and Σx² = 55.

  2. Compute the slope

    m = (nΣxy minus ΣxΣy) / (nΣx² minus (Σx)²). Here that is (5 × 66 minus 15 × 20) / (5 × 55 minus 15²) = 30 / 50 = 0.6. Keep it as a fraction if it does not come out as a short decimal.

  3. Compute the intercept

    b = (Σy minus mΣx) / n. Here (20 minus 0.6 × 15) / 5 = 11/5 = 2.2. Use the exact slope from the previous step, not a rounded one.

  4. Write the line and check it

    y = 0.6x + 2.2. Check that it passes through the point of means (3, 4): 0.6 × 3 + 2.2 = 4. If it does not, one of the sums is wrong.

  5. Judge the fit

    Work out r, or look at the residuals. Here r = √15/5, about 0.7746, and the residuals are -0.8, 0.6, 1, -0.6 and -0.2, with squares adding up to 2.4.

Lines of best fit for small data sets

Each row is one data set typed into the calculator above. The slope and intercept are exact; r is shown as a radical with its decimal.

x valuesy valuesLine of best fitrr²
1, 2, 3, 4, 52, 4, 5, 4, 5y = 0.6x + 2.2√15/5 ≈ 0.77460.6
0, 1, 2, 37, 5, 3, 1y = -2x + 7-11
2, 4, 6, 8, 103, 7, 8, 12, 15y = 1.45x + 0.329√215/430 ≈ 0.9889841/860
1, 2, 3, 4, 5, 652, 58, 61, 67, 70, 76y = (162/35)x + 47.89√15/35 ≈ 0.9959243/245
1, 2, 3, 4, 53, 1, 4, 1, 5y = 0.4x + 1.6√2/4 ≈ 0.35360.125
1, 32, 8y = 3x - 111
-2, -1, 0, 1, 24, 1, 0, 1, 4y = 200

Worked examples

The textbook set

plain
x: 1, 2, 3, 4, 5; y: 2, 4, 5, 4, 5

The sums are Σx = 15, Σy = 20, Σxy = 66 and Σx² = 55 with n = 5. The slope is 30/50 = 0.6 and the intercept is 11/5 = 2.2, so the line is y = 0.6x + 2.2. The residuals are -0.8, 0.6, 1, -0.6 and -0.2, which add to 0, and their squares add to 2.4. The correlation is r = √15/5, about 0.7746: strong and positive.

A slope that is a fraction: hours studied against test score

plain
x: 1, 2, 3, 4, 5, 6; y: 52, 58, 61, 67, 70, 76

The means are 3.5 hours and 64 points. The slope is 162/35, about 4.63 points per hour, and the intercept is 64 minus (162/35) × 3.5 = 47.8, so y = (162/35)x + 47.8 with r ≈ 0.9959. At 4.5 hours the line predicts 2402/35, about 68.63. At 12 hours it predicts about 103.34 on a test marked out of 100: an extrapolation the data cannot support.

Points already on a line

plain
x: 0, 1, 2, 3; y: 7, 5, 3, 1

Every step of 1 in x drops y by 2, so the points lie exactly on y = -2x + 7. The calculator finds that line, every residual is 0, the sum of squared residuals is 0, and r = -1: a perfect negative correlation.

Predicting inside and outside the data

plain
x: 2, 4, 6, 8, 10; y: 3, 7, 8, 12, 15

The line is y = 1.45x + 0.3 with r ≈ 0.9889, and it passes through the point of means (6, 9). At x = 7, inside the data, it predicts 10.45, and that is an interpolation you can trust about as far as the fit. At x = 60 it predicts 87.3, fifty units past the last point, where nothing in the data says the trend still holds.

A strong r that hides a curve

plain
x: 0, 1, 2, 3, 4, 5; y: 0, 1, 4, 9, 16, 25

The line of best fit is y = 5x - 10/3 with r ≈ 0.9599, which sounds excellent. The residuals are 10/3, -2/3, -8/3, -8/3, -2/3 and 10/3: positive at both ends, negative in the middle. That U shape means the points lie on a curve, and quadratic regression fits them exactly as y = x².

Common mistakes

  • Rounding the slope before working out the intercept. With a slope of 162/35, using 4.63 gives an intercept of 47.795 instead of 47.8. Carry the exact slope through, round once at the end.
  • Mixing up nΣx² and (Σx)². For x = 1 to 5 the first is 5 × 55 = 275 and the second is 15² = 225. Swapping them changes the sign of the bottom line.
  • Pairing the values out of order. The first x belongs to the first y. Sorting one column without the other produces a line through points that do not exist.
  • Extrapolating far beyond the data. The line describes the range it was fitted on; a prediction well outside that range can be wildly wrong even when r is close to 1.
  • Taking a high r as proof that a straight line is the right model. The squares 0 to 25 give r ≈ 0.96 and still lie on a curve. Look at the residuals.
  • Reading the slope as a cause. A line says how y tends to change with x in this data, not that changing x would change y.
  • Fitting x on y when the question asks for y on x. The two regression lines are different unless every point lies on one line.

Linear regression FAQ

How do I calculate linear regression by hand?
Make a table with x, y, xy and x² for each point and add the columns. The slope is m = (nΣxy minus ΣxΣy) / (nΣx² minus (Σx)²) and the intercept is b = (Σy minus mΣx) / n. For x = 1 to 5 and y = 2, 4, 5, 4, 5 that gives m = 30/50 = 0.6 and b = 11/5 = 2.2, so y = 0.6x + 2.2. This page prints the same table and fills in both formulas so you can check each line.
What is the least squares regression line?
The straight line that makes the sum of the squared residuals as small as possible, where a residual is a point's actual y minus the y the line predicts. It is unique whenever the x values are not all equal. For the five points above no other line gets the sum of squares below 2.4.
What do the slope and the intercept mean?
The slope is how much y changes, on average, for each increase of 1 in x, in y units per x unit. The intercept is the predicted y when x is 0. The intercept only means something if x = 0 is inside or near the data; a regression of weight on height has an intercept at height 0 that describes no real person.
Is the line of best fit the same as the regression line?
Yes. In statistics the line of best fit is the least squares regression line, which is what this calculator computes. A line of best fit drawn by eye on paper is an estimate of it, and it should pass through the point of means, as the calculated line always does.
What is a good r squared for a regression line?
It depends on the field. r² is the share of the variation in y that the line accounts for, so 0.9 means 90%. Measurements from a controlled physics experiment often reach 0.99; data about people often stops at 0.3 to 0.5 and is still useful. A high r² does not prove a straight line is the right shape, so check the residuals too.
How many points do I need for linear regression?
Two points define a line, so the calculator accepts two, but a line through two points always fits perfectly and says nothing about the data. With three or more points the residuals, r and r² begin to carry information, and more points make the slope more reliable.
Can I use the regression line to predict?
Inside the range of your x values, yes: that is interpolation, and it is as good as the fit. Outside the range it is extrapolation, and nothing in the data says the trend continues. The test-score example on this page predicts about 103 points for 12 hours of study, on a test out of 100.
Why are the answers fractions instead of decimals?
Because the slope and intercept of a least squares line on rational data are rational numbers, and the fraction is the exact value. 162/35 is the slope; 4.6286 is a rounding of it, and every prediction made from the rounding carries the error. The Copy button gives decimals when you need them.

More math tools

Coddy programming languages illustration

Learn math with Coddy

GET STARTED