← Mathematics for Machine LearningDownload PDF
Cheat sheetPrint-friendlyIntermediate beginner

Mathematics for Machine Learning — Student Cheat Sheet

Use this as a one-sitting review before the lab or an interview. Full explanations live in the pillar tutorial. Needs: high-school maths + basic Python/NumPy.

Download the PDFLab ZIP

Symbols in 60 seconds#

Symbol Say it Meaning in ML
\(\mathbf{x}\) "vector x" Features for one example
\(\mathbf{w}\), \(\boldsymbol{\theta}\) "weights" / "theta" Parameters you learn
\(b\) "bias" Offset term
\(\hat{y}\) "y-hat" Prediction
\(J(\mathbf{w})\) "cost J of w" Scalar loss / cost
\(\frac{dJ}{dw}\) "dJ dw" Derivative (1D slope)
\(\frac{\partial J}{\partial w_i}\) "partial J partial w_i" Slope when only \(w_i\) moves
\(\nabla J\) "grad J" / "nabla J" Vector of all partials
\(\eta\) "eta" Learning rate
\(\|\mathbf{a}\|_2\) "L2 norm of a" Euclidean length
\(\mathbf{a}\cdot\mathbf{b}\) "a dot b" Sum of \(a_i b_i\)
\(\mathbf{X}^\top\) "X transpose" Rows ↔ columns

Must-know equations#

Linear prediction

\[\hat{y} = \mathbf{w}\cdot\mathbf{x} + b = \sum_i w_i x_i + b\]

L2 norm / distance

\[\|\mathbf{a}\|_2 = \sqrt{\sum_i a_i^2}, \qquad d(\mathbf{a},\mathbf{b})=\|\mathbf{a}-\mathbf{b}\|_2\]

Cosine similarity

\[\cos(\mathbf{u},\mathbf{v})=\frac{\mathbf{u}\cdot\mathbf{v}}{\|\mathbf{u}\|_2\|\mathbf{v}\|_2}\]

Gradient descent update

\[\mathbf{w} \leftarrow \mathbf{w} - \eta\,\nabla J(\mathbf{w})\]

MSE (batch)

\[J(\boldsymbol{\theta})=\frac{1}{n}\|\mathbf{X}\boldsymbol{\theta}-\mathbf{y}\|_2^2, \qquad \nabla_{\boldsymbol{\theta}}J=\frac{2}{n}\mathbf{X}^\top(\mathbf{X}\boldsymbol{\theta}-\mathbf{y})\]

Confusions to kill early#

Pair Remember
Scalar vs vector Loss \(J\) is a scalar; features/weights are vectors
Derivative vs gradient 1 variable → derivative; many → gradient (a vector)
Dot vs * in NumPy Dot → one number; * → element-wise array
Subtract vs add gradient Subtract to minimise cost
L1 vs L2 L1 = sum of abs; L2 = sqrt of sum of squares

Lab checkpoints (expected numbers)#

Run: python maths_for_ml_lab.py (from the lab ZIP)

Check Expected
\(8\cdot4+3\cdot7+20\) \(73.0\)
Batch preds \([73,\ 83,\ 60]\)
\(\|(3,4)\|_2\) \(5.0\)
Cosine of parallel vectors \(1.0000\)
1D GD final \(w\) (25 steps, \(\eta=0.1\)) \(\approx 2.981\)
2D GD final \(\approx [1.970,\ -1.000]\)
Linear reg recovered \(\boldsymbol{\theta}\) \(\approx [4.99,\ 2.00,\ 0.93]\)

Danger demo: \(\eta=1.1\) on \(J=(w-3)^2\) from \(w=-2\) → cost climbs (\(25,36,52,\ldots\)).

5-line NumPy patterns#

import numpy as np
y_hat = X @ w + b                    # batch predictions
loss = np.mean((y_hat - y) ** 2)     # MSE
grad = (2 / n) * (X.T @ (y_hat - y)) # MSE gradient w.r.t. w (X includes bias col)
w = w - lr * grad                    # descent step
np.linalg.norm(a, ord=2)             # L2 norm

Interview pocket answers#

  1. Gradient vs derivative: gradient = all partials as a vector; derivative = 1D case.
  2. Why minus \(\nabla J\)? Gradient points uphill on cost; we go downhill.
  3. Too-large \(\eta\): overshoot / diverge — loss can increase.
  4. Bias trick: append a column of ones; bias becomes one more weight.
  5. Cosine vs dot: cosine = normalised dot; equal to dot on unit vectors (RAG retrieval).

Next pages on pushpjeet.com#

Newsletter

Liked this? Get the next tutorial by email

New tutorials, labs and cheat sheets, plus AWS batch dates and Oracle tips if you want them. No spam, unsubscribe anytime.

I’m interested in

Double opt-in: you’ll get a confirmation email first. Privacy policy