← Linear Regression BeginnerTry the quiz
Cheat sheetPrint-friendlyBeginner

Linear Regression for Beginners — Student Cheat Sheet

Use this as a one-sitting review before the lab or an interview. Full explanations live in the pillar tutorial.


Symbols in 60 seconds#

Symbol Say it Meaning
\(x\) "x" Independent variable (input you know)
\(y\) "y" Dependent variable (output / target)
\(\hat{y}\) "y-hat" Predicted output
\(m\) "slope" Change in predicted \(y\) when \(x\) increases by 1
\(b\) "intercept" Predicted \(y\) when \(x = 0\)
\(n\) "n" Number of data points
\(\bar{x}, \bar{y}\) "x-bar, y-bar" Means of \(x\) and \(y\)
residual "residual" \(y_{\text{actual}} - y_{\text{predicted}}\)
\(R^2\) "R-squared" Fraction of \(y\) variation explained on this data
RMSE "R-M-S-E" Typical error size, same units as \(y\)

Must-know equations#

Line / prediction

\[ y = mx + b, \qquad \hat{y} = m \cdot x_{\text{new}} + b \]

Fit (simple linear regression)

\[ \bar{x} = \frac{\sum x}{n}, \qquad \bar{y} = \frac{\sum y}{n} \]
\[ m = \frac{\sum (x - \bar{x})(y - \bar{y})}{\sum (x - \bar{x})^2}, \qquad b = \bar{y} - m \cdot \bar{x} \]

Error metrics (this course uses \(\mathrm{RMSE} = \sqrt{\mathrm{SS_{res}}/n}\))

\[ \mathrm{SS_{res}} = \sum (\text{residual})^2, \qquad \mathrm{SS_{tot}} = \sum (y - \bar{y})^2 \]
\[ R^2 = 1 - \frac{\mathrm{SS_{res}}}{\mathrm{SS_{tot}}}, \qquad \mathrm{RMSE} = \sqrt{\frac{\mathrm{SS_{res}}}{n}} \]

When to use / when not#

Use simple LR when… Prefer something else when…
\(y\) is a continuous number Target is a category (spam / not spam)
Relationship looks roughly linear Clear curve / strong non-linearity
You need a baseline, explainable model Complex interactions and enough data for them
One main numeric driver Inputs are text/images without numeric features

Verified reference numbers (do not invent others)#

Dataset Line \(R^2\) Extra check
Houses: \(x=[1,1.5,2,2.5,3]\), \(y=[150,200,240,300,330]\) \(y = 92x + 60\) ≈ 0.9925 pred@2.2 → 262.4
Marks: \(x=[2,3,5,7,8]\), \(y=[40,50,65,80,85]\) \(y = 7.5x + 26.5\) ≈ 0.9949 pred@6 → 71.5
Lab 1: \((1,3),(2,5),(3,7),(4,10)\) \(y = 2.3x + 0.5\) — pred@5 → 12.0

Confusions to kill early#

Pair Remember
\(x\) vs \(y\) \(x\) = input you know; \(y\) = what you predict
Slope vs intercept \(m\) = change per +1 \(x\); \(b\) = value at \(x=0\)
Residual vs error metric Residual is per point; RMSE / \(R^2\) summarize the set
\(R^2\) vs RMSE \(R^2\) near 1 = strong fit on data; RMSE = error size in \(y\) units
High \(R^2\) vs "true model" Fit on training data ≠ proof of cause or future accuracy
Forgetting \(b\) \(7.5\times 6 = 45\) is wrong; need \(+ 26.5 \rightarrow 71.5\)

Lab checkpoints#

Lab 1 (paper): \(\bar{x}=2.5\), \(\bar{y}=6.25\), \(m=2.3\), \(b=0.5\), residuals sum to 0, RMSE ≈ 0.274.

Lab 2 (Python CSV): \(m=7.5\), \(b=26.5\), \(R^2=0.994898\), RMSE=1.224745, pred@6=71.5.


Interview anchors (one line each)#

  1. What is LR? Fit \(y=mx+b\) with least squares to predict a numeric \(y\).
  2. Least squares? Minimize sum of squared (actual − predicted).
  3. Good fit? Low RMSE, \(R^2\) toward 1, residuals without a leftover pattern.
  4. Simple vs multiple? One input vs several inputs; same goal, more coefficients.
  5. Cloud example shape? Cost vs VM count, or latency vs load — only if the scatter looks roughly linear.

Next reading on pushpjeet.com#