← Linear Regression BeginnerCheat sheet
Interactive10 questionsNo install

Linear Regression for Beginners — Interactive Quiz

Linear Regression — Beginner

Training pack for Cloud / DevOps learners · Pushpjeet Cholkar

Straight-line prediction, from zero

Plain words. Concrete numbers. No analogies. Built for students who may have never taken a statistics course.

Beginner Least squares R² & RMSE Pen + Python labs Interactive quiz

What it is

Linear regression is a method that finds a straight line that best describes how one number changes when another number changes.

You give it pairs of numbers (input, output). It returns a line equation you can use to predict the output for a new input.

When we use it

  1. You want to predict a continuous number (price, marks, response time, cost).
  2. You believe the relationship is roughly a straight line.
  3. You have labeled examples — past pairs of (input, known output).
GoalInput (x)Output (y)
Predict house priceSize (sq ft)Price
Predict exam marksStudy hoursMarks
Predict cloud billNumber of VMsMonthly cost
Predict latencyRequest size (KB)Response time (ms)

The line equation: y = mx + b

y = m·x + b
SymbolNamePlain meaning
xIndependent variableThe input you already know
yDependent variableThe output you want to predict
mSlopeHow much y changes when x increases by 1
bInterceptThe value of y when x is 0

Prediction for a new input: predicted_y = m * x + b

How we find the best line (least squares)

  1. For each data point, the line gives a predicted y.
  2. Residual = actual y − predicted y.
  3. Square each residual.
  4. Add all squared residuals.
  5. Choose m and b that make this sum as small as possible.
mean_x = Σx / n
mean_y = Σy / n
m = Σ((x − mean_x)(y − mean_y)) / Σ((x − mean_x)²)
b = mean_y − m · mean_x

Residuals

residual = actual_y − predicted_y
  • Residual > 0 → line predicted too low
  • Residual < 0 → line predicted too high
  • Residual = 0 → exact hit

Simple vs multiple (briefly)

TypeInputsEquation
SimpleOne xy = m·x + b
MultipleSeveral x₁, x₂, …y = m₁x₁ + m₂x₂ + … + b

This pack focuses on simple linear regression.

Assumptions (stated simply)

  1. Linearity — relationship is close to a straight line.
  2. Independence — points are not strongly chained in a way you ignore.
  3. Homoscedasticity — residual spread is roughly similar across x.
  4. No extreme single-point control — check outliers.
  5. Errors look roughly random — residuals should not show a clear leftover pattern.

Judging fit: R² and RMSE

RMSE = sqrt( (Σ residual²) / n )     # lower better; same units as y
R²   = 1 − SS_res / SS_tot           # closer to 1 = stronger fit on that data

R² near 1 means the line explains most of the variation in y on the data you used. It does not by itself prove the model will work on new data.

Key terms (1–2 sentences each)

Linear regressionA method that fits a straight-line equation to data so you can predict a numeric output from one or more numeric inputs.
Dependent variable (y)The number you are trying to predict or explain.
Independent variable (x)The number you use as input to make the prediction.
Intercept (b)The predicted value of y when every input is zero.
Slope (m)How much the predicted y changes when the input x increases by one unit.
ResidualActual y minus predicted y for one data point.
Least squaresThe rule that chooses the line by minimizing the sum of squared residuals.
R-squared (R²)A number (usually between 0 and 1) that says what fraction of the variation in y is explained by the line.
Overfitting (light)Making a model that follows the training data too tightly, including noise, so it predicts poorly on new data.

Example 1 — House price from size

x = size in thousands of sq ft  ·  y = price in $1000s

xy
1.0150
1.5200
2.0240
2.5300
3.0330
mean_x = 2.0
mean_y = 244.0
Σ(x−mean_x)(y−mean_y) = 230.0
Σ(x−mean_x)² = 2.5
m = 230 / 2.5 = 92.0
b = 244 − 92×2 = 60.0

Line: y = 92·x + 60
xactualpredictedresidual
1.0150152−2
1.5200198+2
2.0240244−4
2.5300290+10
3.0330336−6

R² ≈ 0.9925  ·  RMSE ≈ 5.6569 ($1000s)

New prediction for size 2.2: 92×2.2 + 60 = 262.4 → about $262,400.

Example 2 — Study hours to marks

x = hours studied  ·  y = marks

xy
240
350
565
780
885
mean_x = 5.0
mean_y = 64.0
Σ(x−mean_x)(y−mean_y) = 195
Σ(x−mean_x)² = 26
m = 195 / 26 = 7.5
b = 64 − 7.5×5 = 26.5

Line: y = 7.5·x + 26.5
xactualpredictedresidual
24041.5−1.5
35049.0+1.0
56564.0+1.0
78079.0+1.0
88586.5−1.5

R² ≈ 0.9949  ·  RMSE ≈ 1.2247 marks

New prediction for 6 hours: 7.5×6 + 26.5 = 71.5.

Lab 1 — Pen and paper

Data: (1,3), (2,5), (3,7), (4,10)

Expected: mean_x = 2.5, mean_y = 6.25, m = 2.3, b = 0.5 → y = 2.3x + 0.5

Predictions: 2.8, 5.1, 7.4, 9.7. Sum of residuals ≈ 0. Prediction for x=5 → 12.0.

Full step table is in 02-labs.md.

Lab 2 — Python (standard library, no cloud cost)

Create CSV, fit with plain Python. Expected: m = 7.5, b = 26.5, R² = 0.994898, RMSE = 1.224745, pred(6)=71.5.

mkdir -p lr-lab2 && cd lr-lab2
cat > study_marks.csv << 'CSV'
hours,marks
2,40
3,50
5,65
7,80
8,85
CSV

Then run the Python block from 02-labs.md (csv + least-squares formulas). Cleanup: rm -rf lr-lab2.

sklearn is not required. numpy is optional (np.polyfit should also return 7.5 and 26.5).

Interactive quiz — 10 questions

Instant feedback per question. Earn XP. Unlock badges. Original teaching questions (not from real exams).

Score
0 / 10
XP
0
Level
Novice
Progress
0%
Starter Line Fitter Residual Reader Metric Mind Regression Rookie

Final results

0 / 10

XP earned: 0 · Level: Novice

Cheat sheet

y = m·x + b
m = Σ((x−mean_x)(y−mean_y)) / Σ((x−mean_x)²)
b = mean_y − m·mean_x
residual = y − ŷ
R² = 1 − SS_res/SS_tot
RMSE = sqrt(SS_res / n)

Use when y is continuous and the pattern looks linear. Prefer another method for category labels or clear curves.

Common mistakes

  1. Swapping x and y
  2. Forgetting the intercept in predictions
  3. Treating high R² as proof the model is “true”
  4. Ignoring residual patterns
  5. Using regression for yes/no classification without a proper method
  6. Comparing RMSE values that used different divisors (n vs n−2) without stating it
  7. Letting one outlier dominate the line

Interview prompts (short model answers)

What is linear regression?

It fits a straight line y = mx + b to numeric data using least squares, then predicts y for new x.

Slope vs intercept?

Slope: change in prediction per +1 input. Intercept: prediction when input is 0.

What is least squares?

Minimize the sum of squared residuals (actual − predicted).

How do you judge fit?

RMSE (lower better), R² (closer to 1 stronger on that data), plus residual checks. Training fit ≠ future performance.

Simple vs multiple?

Simple: one input. Multiple: several inputs, same squared-error goal, more coefficients to estimate.

Linear Regression — Beginner teaching pack · Standalone HTML (no CDN) · Pushpjeet Cholkar training materials