Symbols in 60 seconds#
| Symbol | Meaning |
|---|---|
| \(n\), \(p\) | Rows; features (not counting the intercept) |
| \(\mathbf{X}\) | Design matrix, \(n \times (p+1)\), first column all ones |
| \(\mathbf{y}\), \(\boldsymbol{\beta}\), \(\boldsymbol{\varepsilon}\) | Target vector; coefficients (intercept first); errors |
| \(\hat{\boldsymbol{\beta}}\), \(\hat{\mathbf{y}}\), \(\mathbf{e}\) | Estimated coefficients; fitted values; residuals \(\mathbf{y} - \hat{\mathbf{y}}\) |
| \(\mathbf{H}\), \(h_{ii}\) | Hat matrix; leverage of row \(i\) |
| \(\hat{\sigma}^2\) | \(\mathrm{SS_{res}}/(n-p-1)\), residual variance estimate |
| \(\kappa\) | Condition number (largest / smallest singular value) |
| \(\eta\) | Learning rate (gradient descent) |
\(\lambda\), alpha |
Penalty strength (Ridge alpha = \(\lambda\); Lasso alpha uses a \(1/(2n)\) loss scaling) |
Must-know equations#
Model and OLS
Compute with QR/SVD (np.linalg.lstsq, LinearRegression), never an explicit inverse: \(\kappa(\mathbf{X}^\top\mathbf{X}) = \kappa(\mathbf{X})^2\).
Always true after an OLS fit: \(\mathbf{X}^\top\mathbf{e} = \mathbf{0}\); with an intercept \(\sum e_i = 0\).
Hat matrix: \(\mathbf{H} = \mathbf{X}(\mathbf{X}^\top\mathbf{X})^{-1}\mathbf{X}^\top\), \(\hat{\mathbf{y}} = \mathbf{H}\mathbf{y}\), \(\mathbf{H}^2 = \mathbf{H}\), \(\operatorname{trace}(\mathbf{H}) = p+1\), average leverage \((p+1)/n\).
Inference
CI for \(\beta_j\): \(\hat{\beta}_j \pm t_{1-\alpha/2,\,n-p-1}\cdot\operatorname{SE}(\hat{\beta}_j)\). Prediction interval = CI for the mean plus \(\sigma^2\) under the square root, so always wider.
Fit metrics
\(R^2\) never falls when you add a feature (in-sample OLS with intercept); adjusted \(R^2\) can. Test-set \(R^2\) can be negative.
Gradient descent (MSE)
Back-transform from standardised features: \(\beta_j = w_j/s_j\), \(\beta_0 = w_0 - \sum_j w_j\mu_j/s_j\).
Regularisation
Ridge: always unique for \(\lambda > 0\), no exact zeros. Lasso: exact zeros, unstable choice among correlated features. Elastic Net: both (l1_ratio). Standardise first; never penalise the intercept.
Diagnostics formulas
Rules of thumb (screening conventions, not hard limits)#
| Check | Look closer when… |
|---|---|
| VIF | > 5 (attention), > 10 (often called serious) |
| Condition number, standardised \(\mathbf{X}\) | > 30 (statsmodels warns > 1000 on the unscaled design; that can be units, not collinearity) |
| Leverage \(h_{ii}\) | > \(2(p+1)/n\) |
| Cook's distance | > 1 large; > \(4/n\) worth a look |
| Durbin–Watson | Outside roughly 1.5–2.5, and only if row order means something (time) |
| Studentised residual | \(\lvert r_i \rvert\) > 2 or 3 |
| Breusch–Pagan | Small p (for example < 0.05): use HC3 robust SEs, transform \(y\), or WLS |
What each assumption failure does#
| Failure | Coefficients | Standard errors | Fix |
|---|---|---|---|
| Heteroscedasticity | Still unbiased | Wrong (classical) | HC3 SEs, log \(y\), WLS |
| Autocorrelation | Still unbiased | Usually too small | Newey–West (HAC), lags, time-series model |
| Multicollinearity | Unbiased, unstable | Inflated by \(\sqrt{\mathrm{VIF}}\) | Drop/combine, centre, Ridge, PCA/PLS |
| Omitted confounder (endogeneity) | Biased | Misleading | Better design or data; diagnostics cannot fully detect it |
| Non-normal errors | Unaffected | Exact small-\(n\) t/F lost | Usually fine with large \(n\) |
Verified reference numbers (do not invent others)#
| Item | Value |
|---|---|
| Example (a) data | \(x_1\) = [1..6] VMs, \(x_2\) = [2,1,3,3,5,4] TB, \(y\) = [5,7,12,13,19,19] ($100s) |
| Example (a) \(\mathbf{X}^\top\mathbf{X}\), det | [[6,21,18],[21,91,74],[18,74,64]], 324 |
| Example (a) \(\mathbf{X}^\top\mathbf{y}\) | [75, 316, 263] |
| Example (a) \(\hat{\boldsymbol{\beta}}\) | [2/3, 13/6, 17/12] ≈ [0.6667, 2.1667, 1.4167] |
| Example (a) \(R^2\) / adjusted | 97/98 ≈ 0.9898 / 289/294 ≈ 0.9830 |
| Example (a) SS_res / SS_tot / RMSE / F | 1.75 / 171.5 / 0.5401 / 145.5 |
| Example (a) prediction at (4, 4) | 15.0 (95% PI [12.1933, 17.8067]) |
| Example (b) VIF | 862.59, 862.50, 1.0034 |
| Lab 1 condition numbers | \(\kappa(\mathbf{X})\) 2.089e+06, \(\kappa(\mathbf{X}^\top\mathbf{X})\) 4.358e+12 |
| Lab 2 | unscaled \(\kappa\) 6.667e+07, stable \(\eta\) < 2.059e-07; scaled \(\kappa\) 1.056, stable \(\eta\) < 0.9745; OLS MSE 245.0873 |
| Lab 3 test \(R^2\) (diabetes, 80/20, random_state 42) | OLS 0.4526 · Ridge 0.4544 · Lasso 0.4613 |
| Lab 3 test RMSE | OLS 53.85 · Ridge 53.77 · Lasso 53.43 (target SD 77.09) |
| Lab 3 split check (random_state 0) | 0.3322 / 0.3348 / 0.3354 |
| Lab 3 diagnostics | VIF s1 55.25, s2 35.76; BP p 0.0345; JB p 0.494; DW 1.794; max Cook's D 0.0277 |
Caveats. Examples (a)–(c) use made-up data; their p-values are illustrations only. In Lab 3 the model gaps are smaller than split-to-split variance, so do not claim Lasso is "better". Diagnostic thresholds are rules of thumb.
Code you will type most#
import numpy as np, statsmodels.api as sm
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LinearRegression, RidgeCV, LassoCV
from statsmodels.stats.outliers_influence import variance_inflation_factor
beta = np.linalg.lstsq(X_with_ones, y, rcond=None)[0] # OLS, no explicit inverse
fit = sm.OLS(y, sm.add_constant(X)).fit() # statsmodels: add the constant yourself
robust = sm.OLS(y, sm.add_constant(X)).fit(cov_type="HC3") # same coefficients, robust SEs
Xc = sm.add_constant(X).values
vif = [variance_inflation_factor(Xc, j) for j in range(1, Xc.shape[1])]
pipe = Pipeline([("scale", StandardScaler()), ("model", RidgeCV(alphas=np.logspace(-3, 3, 61)))]) # leakage-safe
Top 8 mistakes#
- Choosing models by training \(R^2\).
sm.OLSwithoutsm.add_constant.- All \(k\) dummies plus an intercept (dummy variable trap).
- Reading a coefficient as causal, or ignoring "holding others constant".
- Believing heteroscedasticity biases coefficients.
- Fitting the scaler on all data before CV (leakage).
- Ridge/Lasso on unscaled features.
- Random splits on time-ordered or grouped data.
15 flashcards#
| # | Question | Answer |
|---|---|---|
| 1 | Normal equations? | \(\mathbf{X}^\top\mathbf{X}\hat{\boldsymbol{\beta}} = \mathbf{X}^\top\mathbf{y}\) |
| 2 | Trace of the hat matrix? | \(p + 1\) |
| 3 | Two properties of OLS residuals (with intercept)? | Sum to zero; orthogonal to every column of \(\mathbf{X}\) |
| 4 | cond(X) vs cond(XᵀX)? | \(\kappa(\mathbf{X}^\top\mathbf{X}) = \kappa(\mathbf{X})^2\) |
| 5 | BLUE? | Best Linear Unbiased Estimator |
| 6 | Is normality a Gauss–Markov assumption? | No; only for exact small-sample t/F |
| 7 | Heteroscedasticity effect on coefficients? | None on bias; breaks classical SEs and efficiency |
| 8 | VIF and SE inflation? | \(1/(1 - R^2_j)\); SE × \(\sqrt{\mathrm{VIF}}\) |
| 9 | DW for no autocorrelation? | About 2 (range 0–4) |
| 10 | Cook's D rules of thumb? | > 1 large; > \(4/n\) worth inspecting |
| 11 | Adjusted \(R^2\)? | \(1 - (1 - R^2)(n-1)/(n-p-1)\) |
| 12 | Ridge closed form? | \((\mathbf{X}^\top\mathbf{X} + \lambda\mathbf{I})^{-1}\mathbf{X}^\top\mathbf{y}\) on centred data |
| 13 | Why does Lasso give exact zeros? | L1 is non-differentiable at 0 → soft-thresholding; L1 ball has corners |
| 14 | Dummy trap fix? | \(k - 1\) dummies with an intercept |
| 15 | CI vs PI? | CI for the mean response; PI for a new observation, wider by \(\sigma^2\) |
Lab checkpoints#
| Lab | You are done when… |
|---|---|
| Example (a) script | beta (lstsq) = [0.6667 2.1667 1.4167], R^2 = 0.9898 adjusted R^2 = 0.9830 |
| Lab 1 | Six routes agree (max diff vs lstsq about 1e-15), trace(H) = 4.0, cond(X_bad^T X_bad) ≈ 4.358e+12 |
| Lab 2 | max GD-vs-OLS difference about 1e-13, scaled final MSE 245.0873, lr=1e-06 unscaled diverges |
| Lab 3 | Test R² table: 0.4526 / 0.4544 / 0.4613; Lasso zero coefficient: ['s2'] |