← Linear Regression ExpertTry the quiz
Cheat sheetPrint-friendlyExpert

Linear Regression with Multiple Regression — Student Cheat Sheet (Expert)

Course companion: Linear Regression with Multiple Regression (Expert) · Instructor: Pushpjeet Cholkar · Level: Advanced / Expert · Needs: matrix basics, Python 3.10+, numpy, scikit-learn, statsmodels

Use this for a one-sitting review before the labs or an interview. Derivations and explanations are in the tutorial. Previous level: Linear Regression for Beginners.


Symbols in 60 seconds#

Symbol Meaning
\(n\), \(p\) Rows; features (not counting the intercept)
\(\mathbf{X}\) Design matrix, \(n \times (p+1)\), first column all ones
\(\mathbf{y}\), \(\boldsymbol{\beta}\), \(\boldsymbol{\varepsilon}\) Target vector; coefficients (intercept first); errors
\(\hat{\boldsymbol{\beta}}\), \(\hat{\mathbf{y}}\), \(\mathbf{e}\) Estimated coefficients; fitted values; residuals \(\mathbf{y} - \hat{\mathbf{y}}\)
\(\mathbf{H}\), \(h_{ii}\) Hat matrix; leverage of row \(i\)
\(\hat{\sigma}^2\) \(\mathrm{SS_{res}}/(n-p-1)\), residual variance estimate
\(\kappa\) Condition number (largest / smallest singular value)
\(\eta\) Learning rate (gradient descent)
\(\lambda\), alpha Penalty strength (Ridge alpha = \(\lambda\); Lasso alpha uses a \(1/(2n)\) loss scaling)

Must-know equations#

Model and OLS

\[ \mathbf{y} = \mathbf{X}\boldsymbol{\beta} + \boldsymbol{\varepsilon}, \qquad \mathbf{X}^\top\mathbf{X}\hat{\boldsymbol{\beta}} = \mathbf{X}^\top\mathbf{y}, \qquad \hat{\boldsymbol{\beta}} = (\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{X}^\top\mathbf{y} \]

Compute with QR/SVD (np.linalg.lstsq, LinearRegression), never an explicit inverse: \(\kappa(\mathbf{X}^\top\mathbf{X}) = \kappa(\mathbf{X})^2\).

Always true after an OLS fit: \(\mathbf{X}^\top\mathbf{e} = \mathbf{0}\); with an intercept \(\sum e_i = 0\).

Hat matrix: \(\mathbf{H} = \mathbf{X}(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{X}^\top\), \(\hat{\mathbf{y}} = \mathbf{H}\mathbf{y}\), \(\mathbf{H}^2 = \mathbf{H}\), \(\operatorname{trace}(\mathbf{H}) = p+1\), average leverage \((p+1)/n\).

Inference

\[ \operatorname{SE}(\hat{\beta}_j) = \hat{\sigma}\sqrt{\big[(\mathbf{X}^\top\mathbf{X})^{-1}\big]_{jj}}, \qquad t_j = \frac{\hat{\beta}_j}{\operatorname{SE}(\hat{\beta}_j)}, \qquad F = \frac{\mathrm{SS_{reg}}/p}{\mathrm{SS_{res}}/(n-p-1)} \]

CI for \(\beta_j\): \(\hat{\beta}_j \pm t_{1-\alpha/2,\,n-p-1}\cdot\operatorname{SE}(\hat{\beta}_j)\). Prediction interval = CI for the mean plus \(\sigma^2\) under the square root, so always wider.

Fit metrics

\[ R^2 = 1 - \frac{\mathrm{SS_{res}}}{\mathrm{SS_{tot}}}, \qquad R^2_{\text{adj}} = 1 - (1 - R^2)\frac{n-1}{n-p-1}, \qquad \mathrm{RMSE} = \sqrt{\mathrm{SS_{res}}/n} \]

\(R^2\) never falls when you add a feature (in-sample OLS with intercept); adjusted \(R^2\) can. Test-set \(R^2\) can be negative.

Gradient descent (MSE)

\[ \boldsymbol{\beta} \leftarrow \boldsymbol{\beta} - \eta\cdot\frac{2}{n}\mathbf{X}^\top(\mathbf{X}\boldsymbol{\beta} - \mathbf{y}), \qquad \text{stable if } \eta < \frac{1}{\lambda_{\max}(\mathbf{X}^\top\mathbf{X}/n)} \]

Back-transform from standardised features: \(\beta_j = w_j/s_j\), \(\beta_0 = w_0 - \sum_j w_j\mu_j/s_j\).

Regularisation

\[ \text{Ridge: } (\mathbf{X}^\top\mathbf{X} + \lambda\mathbf{I})^{-1}\mathbf{X}^\top\mathbf{y}, \qquad \text{Lasso (orthonormal design): } \operatorname{sign}(z)\max(\lvert z \rvert - \lambda, 0) \]

Ridge: always unique for \(\lambda > 0\), no exact zeros. Lasso: exact zeros, unstable choice among correlated features. Elastic Net: both (l1_ratio). Standardise first; never penalise the intercept.

Diagnostics formulas

\[ \mathrm{VIF}_j = \frac{1}{1 - R^2_j}, \qquad D_i = \frac{e_i^2}{(p+1)\hat{\sigma}^2}\cdot\frac{h_{ii}}{(1-h_{ii})^2}, \qquad \mathrm{DW} \approx 2(1 - \hat{\rho}_1) \]

Rules of thumb (screening conventions, not hard limits)#

Check Look closer when…
VIF > 5 (attention), > 10 (often called serious)
Condition number, standardised \(\mathbf{X}\) > 30 (statsmodels warns > 1000 on the unscaled design; that can be units, not collinearity)
Leverage \(h_{ii}\) > \(2(p+1)/n\)
Cook's distance > 1 large; > \(4/n\) worth a look
Durbin–Watson Outside roughly 1.5–2.5, and only if row order means something (time)
Studentised residual \(\lvert r_i \rvert\) > 2 or 3
Breusch–Pagan Small p (for example < 0.05): use HC3 robust SEs, transform \(y\), or WLS

What each assumption failure does#

Failure Coefficients Standard errors Fix
Heteroscedasticity Still unbiased Wrong (classical) HC3 SEs, log \(y\), WLS
Autocorrelation Still unbiased Usually too small Newey–West (HAC), lags, time-series model
Multicollinearity Unbiased, unstable Inflated by \(\sqrt{\mathrm{VIF}}\) Drop/combine, centre, Ridge, PCA/PLS
Omitted confounder (endogeneity) Biased Misleading Better design or data; diagnostics cannot fully detect it
Non-normal errors Unaffected Exact small-\(n\) t/F lost Usually fine with large \(n\)

Verified reference numbers (do not invent others)#

Item Value
Example (a) data \(x_1\) = [1..6] VMs, \(x_2\) = [2,1,3,3,5,4] TB, \(y\) = [5,7,12,13,19,19] ($100s)
Example (a) \(\mathbf{X}^\top\mathbf{X}\), det [[6,21,18],[21,91,74],[18,74,64]], 324
Example (a) \(\mathbf{X}^\top\mathbf{y}\) [75, 316, 263]
Example (a) \(\hat{\boldsymbol{\beta}}\) [2/3, 13/6, 17/12] ≈ [0.6667, 2.1667, 1.4167]
Example (a) \(R^2\) / adjusted 97/98 ≈ 0.9898 / 289/294 ≈ 0.9830
Example (a) SS_res / SS_tot / RMSE / F 1.75 / 171.5 / 0.5401 / 145.5
Example (a) prediction at (4, 4) 15.0 (95% PI [12.1933, 17.8067])
Example (b) VIF 862.59, 862.50, 1.0034
Lab 1 condition numbers \(\kappa(\mathbf{X})\) 2.089e+06, \(\kappa(\mathbf{X}^\top\mathbf{X})\) 4.358e+12
Lab 2 unscaled \(\kappa\) 6.667e+07, stable \(\eta\) < 2.059e-07; scaled \(\kappa\) 1.056, stable \(\eta\) < 0.9745; OLS MSE 245.0873
Lab 3 test \(R^2\) (diabetes, 80/20, random_state 42) OLS 0.4526 · Ridge 0.4544 · Lasso 0.4613
Lab 3 test RMSE OLS 53.85 · Ridge 53.77 · Lasso 53.43 (target SD 77.09)
Lab 3 split check (random_state 0) 0.3322 / 0.3348 / 0.3354
Lab 3 diagnostics VIF s1 55.25, s2 35.76; BP p 0.0345; JB p 0.494; DW 1.794; max Cook's D 0.0277

Caveats. Examples (a)–(c) use made-up data; their p-values are illustrations only. In Lab 3 the model gaps are smaller than split-to-split variance, so do not claim Lasso is "better". Diagnostic thresholds are rules of thumb.


Code you will type most#

import numpy as np, statsmodels.api as sm
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LinearRegression, RidgeCV, LassoCV
from statsmodels.stats.outliers_influence import variance_inflation_factor

beta = np.linalg.lstsq(X_with_ones, y, rcond=None)[0]           # OLS, no explicit inverse
fit = sm.OLS(y, sm.add_constant(X)).fit()                        # statsmodels: add the constant yourself
robust = sm.OLS(y, sm.add_constant(X)).fit(cov_type="HC3")       # same coefficients, robust SEs
Xc = sm.add_constant(X).values
vif = [variance_inflation_factor(Xc, j) for j in range(1, Xc.shape[1])]
pipe = Pipeline([("scale", StandardScaler()), ("model", RidgeCV(alphas=np.logspace(-3, 3, 61)))])  # leakage-safe

Top 8 mistakes#

  1. Choosing models by training \(R^2\).
  2. sm.OLS without sm.add_constant.
  3. All \(k\) dummies plus an intercept (dummy variable trap).
  4. Reading a coefficient as causal, or ignoring "holding others constant".
  5. Believing heteroscedasticity biases coefficients.
  6. Fitting the scaler on all data before CV (leakage).
  7. Ridge/Lasso on unscaled features.
  8. Random splits on time-ordered or grouped data.

15 flashcards#

# Question Answer
1 Normal equations? \(\mathbf{X}^\top\mathbf{X}\hat{\boldsymbol{\beta}} = \mathbf{X}^\top\mathbf{y}\)
2 Trace of the hat matrix? \(p + 1\)
3 Two properties of OLS residuals (with intercept)? Sum to zero; orthogonal to every column of \(\mathbf{X}\)
4 cond(X) vs cond(XᵀX)? \(\kappa(\mathbf{X}^\top\mathbf{X}) = \kappa(\mathbf{X})^2\)
5 BLUE? Best Linear Unbiased Estimator
6 Is normality a Gauss–Markov assumption? No; only for exact small-sample t/F
7 Heteroscedasticity effect on coefficients? None on bias; breaks classical SEs and efficiency
8 VIF and SE inflation? \(1/(1 - R^2_j)\); SE × \(\sqrt{\mathrm{VIF}}\)
9 DW for no autocorrelation? About 2 (range 0–4)
10 Cook's D rules of thumb? > 1 large; > \(4/n\) worth inspecting
11 Adjusted \(R^2\)? \(1 - (1 - R^2)(n-1)/(n-p-1)\)
12 Ridge closed form? \((\mathbf{X}^\top\mathbf{X} + \lambda\mathbf{I})^{-1}\mathbf{X}^\top\mathbf{y}\) on centred data
13 Why does Lasso give exact zeros? L1 is non-differentiable at 0 → soft-thresholding; L1 ball has corners
14 Dummy trap fix? \(k - 1\) dummies with an intercept
15 CI vs PI? CI for the mean response; PI for a new observation, wider by \(\sigma^2\)

Lab checkpoints#

Lab You are done when…
Example (a) script beta (lstsq) = [0.6667 2.1667 1.4167], R^2 = 0.9898 adjusted R^2 = 0.9830
Lab 1 Six routes agree (max diff vs lstsq about 1e-15), trace(H) = 4.0, cond(X_bad^T X_bad) ≈ 4.358e+12
Lab 2 max GD-vs-OLS difference about 1e-13, scaled final MSE 245.0873, lr=1e-06 unscaled diverges
Lab 3 Test R² table: 0.4526 / 0.4544 / 0.4613; Lasso zero coefficient: ['s2']