Linear Regression — Beginner
Training pack for Cloud / DevOps learners · Pushpjeet Cholkar
Straight-line prediction, from zero
Plain words. Concrete numbers. No analogies. Built for students who may have never taken a statistics course.
What it is
Linear regression is a method that finds a straight line that best describes how one number changes when another number changes.
You give it pairs of numbers (input, output). It returns a line equation you can use to predict the output for a new input.
When we use it
- You want to predict a continuous number (price, marks, response time, cost).
- You believe the relationship is roughly a straight line.
- You have labeled examples — past pairs of (input, known output).
| Goal | Input (x) | Output (y) |
|---|---|---|
| Predict house price | Size (sq ft) | Price |
| Predict exam marks | Study hours | Marks |
| Predict cloud bill | Number of VMs | Monthly cost |
| Predict latency | Request size (KB) | Response time (ms) |
The line equation: y = mx + b
y = m·x + b
| Symbol | Name | Plain meaning |
|---|---|---|
x | Independent variable | The input you already know |
y | Dependent variable | The output you want to predict |
m | Slope | How much y changes when x increases by 1 |
b | Intercept | The value of y when x is 0 |
Prediction for a new input: predicted_y = m * x + b
How we find the best line (least squares)
- For each data point, the line gives a predicted
y. - Residual = actual
y− predictedy. - Square each residual.
- Add all squared residuals.
- Choose
mandbthat make this sum as small as possible.
mean_x = Σx / n mean_y = Σy / n m = Σ((x − mean_x)(y − mean_y)) / Σ((x − mean_x)²) b = mean_y − m · mean_x
Residuals
residual = actual_y − predicted_y
- Residual > 0 → line predicted too low
- Residual < 0 → line predicted too high
- Residual = 0 → exact hit
Simple vs multiple (briefly)
| Type | Inputs | Equation |
|---|---|---|
| Simple | One x | y = m·x + b |
| Multiple | Several x₁, x₂, … | y = m₁x₁ + m₂x₂ + … + b |
This pack focuses on simple linear regression.
Assumptions (stated simply)
- Linearity — relationship is close to a straight line.
- Independence — points are not strongly chained in a way you ignore.
- Homoscedasticity — residual spread is roughly similar across
x. - No extreme single-point control — check outliers.
- Errors look roughly random — residuals should not show a clear leftover pattern.
Judging fit: R² and RMSE
RMSE = sqrt( (Σ residual²) / n ) # lower better; same units as y R² = 1 − SS_res / SS_tot # closer to 1 = stronger fit on that data
R² near 1 means the line explains most of the variation in y on the data you used. It does not by itself prove the model will work on new data.
Key terms (1–2 sentences each)
Example 1 — House price from size
x = size in thousands of sq ft · y = price in $1000s
| x | y |
|---|---|
| 1.0 | 150 |
| 1.5 | 200 |
| 2.0 | 240 |
| 2.5 | 300 |
| 3.0 | 330 |
mean_x = 2.0 mean_y = 244.0 Σ(x−mean_x)(y−mean_y) = 230.0 Σ(x−mean_x)² = 2.5 m = 230 / 2.5 = 92.0 b = 244 − 92×2 = 60.0 Line: y = 92·x + 60
| x | actual | predicted | residual |
|---|---|---|---|
| 1.0 | 150 | 152 | −2 |
| 1.5 | 200 | 198 | +2 |
| 2.0 | 240 | 244 | −4 |
| 2.5 | 300 | 290 | +10 |
| 3.0 | 330 | 336 | −6 |
R² ≈ 0.9925 · RMSE ≈ 5.6569 ($1000s)
New prediction for size 2.2: 92×2.2 + 60 = 262.4 → about $262,400.
Example 2 — Study hours to marks
x = hours studied · y = marks
| x | y |
|---|---|
| 2 | 40 |
| 3 | 50 |
| 5 | 65 |
| 7 | 80 |
| 8 | 85 |
mean_x = 5.0 mean_y = 64.0 Σ(x−mean_x)(y−mean_y) = 195 Σ(x−mean_x)² = 26 m = 195 / 26 = 7.5 b = 64 − 7.5×5 = 26.5 Line: y = 7.5·x + 26.5
| x | actual | predicted | residual |
|---|---|---|---|
| 2 | 40 | 41.5 | −1.5 |
| 3 | 50 | 49.0 | +1.0 |
| 5 | 65 | 64.0 | +1.0 |
| 7 | 80 | 79.0 | +1.0 |
| 8 | 85 | 86.5 | −1.5 |
R² ≈ 0.9949 · RMSE ≈ 1.2247 marks
New prediction for 6 hours: 7.5×6 + 26.5 = 71.5.
Lab 1 — Pen and paper
Data: (1,3), (2,5), (3,7), (4,10)
Expected: mean_x = 2.5, mean_y = 6.25, m = 2.3, b = 0.5 → y = 2.3x + 0.5
Predictions: 2.8, 5.1, 7.4, 9.7. Sum of residuals ≈ 0. Prediction for x=5 → 12.0.
Full step table is in 02-labs.md.
Lab 2 — Python (standard library, no cloud cost)
Create CSV, fit with plain Python. Expected: m = 7.5, b = 26.5, R² = 0.994898, RMSE = 1.224745, pred(6)=71.5.
mkdir -p lr-lab2 && cd lr-lab2 cat > study_marks.csv << 'CSV' hours,marks 2,40 3,50 5,65 7,80 8,85 CSV
Then run the Python block from 02-labs.md (csv + least-squares formulas). Cleanup: rm -rf lr-lab2.
sklearn is not required. numpy is optional (np.polyfit should also return 7.5 and 26.5).
Interactive quiz — 10 questions
Instant feedback per question. Earn XP. Unlock badges. Original teaching questions (not from real exams).
Final results
XP earned: 0 · Level: Novice
Cheat sheet
y = m·x + b m = Σ((x−mean_x)(y−mean_y)) / Σ((x−mean_x)²) b = mean_y − m·mean_x residual = y − ŷ R² = 1 − SS_res/SS_tot RMSE = sqrt(SS_res / n)
Use when y is continuous and the pattern looks linear. Prefer another method for category labels or clear curves.
Common mistakes
- Swapping x and y
- Forgetting the intercept in predictions
- Treating high R² as proof the model is “true”
- Ignoring residual patterns
- Using regression for yes/no classification without a proper method
- Comparing RMSE values that used different divisors (n vs n−2) without stating it
- Letting one outlier dominate the line
Interview prompts (short model answers)
What is linear regression?
It fits a straight line y = mx + b to numeric data using least squares, then predicts y for new x.
Slope vs intercept?
Slope: change in prediction per +1 input. Intercept: prediction when input is 0.
What is least squares?
Minimize the sum of squared residuals (actual − predicted).
How do you judge fit?
RMSE (lower better), R² (closer to 1 stronger on that data), plus residual checks. Training fit ≠ future performance.
Simple vs multiple?
Simple: one input. Multiple: several inputs, same squared-error goal, more coefficients to estimate.