Relationships: correlation and regression
Measure how two variables move together, fit a line by least squares, and know what neither can tell you.
- Compute and interpret the correlation coefficient r
- Fit a least-squares line and use it to predict
- Explain why correlation doesn’t imply causation
The Pearson correlation r measures how closely two numeric variables follow a straight line: +1 is a perfect upward line, −1 a perfect downward line, 0 no linear relationship. Rough guide: |r| ≥ 0.7 strong, 0.3-0.7 moderate, below 0.3 weak.
Linear regression goes further and fits the line y = slope × x + intercept that makes the sum of squared errors (the vertical gaps between points and line, squared) as small as possible - “least squares”. The slope says how much y changes, on average, per unit of x.
Try it
Fit the line yourself
Move the sliders to make the red residual lines as short as possible overall, and get within 10% of the best possible error. Then reveal the least-squares line and compare.
Your sum of squared errors: 1118.0
1import statistics
2hours = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
3scores = [52, 55, 61, 60, 68, 70, 75, 74, 82, 85]
4fit = statistics.linear_regression(hours, scores)
5print(round(statistics.correlation(hours, scores), 3))
6print(f"score = {fit.slope:.2f} x hours + {fit.intercept:.2f}")0.987 score = 3.62 x hours + 48.27
Correlation is not causation
Ice-cream sales and drownings rise together - because both rise in summer. A third variable that drives both is a confounder. Other traps:
- Reverse causation - does exercise improve mood, or do happier people exercise more?
- Extrapolation - the line fits 1-10 hours of study; it doesn’t promise 200% at 40 hours.
- Non-linear patterns - r only measures straight-line relationships (remember Anscombe’s curved dataset II).
To claim causation you need a design that rules these out - ideally a randomized experiment, which you’ll meet in the A/B testing lesson.
Key takeaways
r measures the strength and direction of a linear relationship, from −1 to +1.
Least squares picks the line with the smallest sum of squared residuals.
Confounders, reverse causation and extrapolation make correlation a poor guide to cause.
Lesson quiz
6 questions · pass with 5 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Practice: write Python
Write Python in the editor and run it against sample inputs. Python runs locally in your browser using a WebAssembly runtime.
Compute r by hand
The first line holds x values, the second y values (space-separated). Compute Pearson’s r without statistics.correlation: sum of (x − x̄)(y − ȳ) divided by the square root of [sum of (x − x̄)² × sum of (y − ȳ)²]. Print r = R (3 decimals) and a description: strong/moderate/weak (|r| ≥ 0.7, ≥ 0.3, otherwise) plus positive/negative.
- Study hours and scores
- Price and sales
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Fit a least-squares line and predict
Each line before the last is a point x y; the last line is predict X. Compute slope = Σ(x − x̄)(y − ȳ) / Σ(x − x̄)² and intercept = ȳ − slope × x̄. Print y = Sx + I (2 decimals) and prediction at X: P (1 decimal).
- Study hours
Python runs in a sandboxed browser worker with a 60 second time limit. Its runtime loads from the Pyodide CDN; your code stays in this browser.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…