Learning objectives#
By the end you will be able to:
- Explain why a prediction is a weighted sum of features (dot product / matrix-vector multiply).
- Distinguish scalar vs vector, row vs column, derivative vs gradient, and L1 vs L2 norms in plain language.
- Compute a dot product, an L2 norm, and a Euclidean distance by hand and in NumPy.
- Write the gradient of a simple cost and explain why it points toward steepest ascent.
- Implement gradient descent from scratch on a 1D and a 2D quadratic cost, and on a tiny linear regression problem.
- Connect that maths to linear regression training and to the first layer of a neural net (preview only).
- Spot common mistakes: learning rate too large, confusing transpose shapes, forgetting the bias term.
Prerequisites#
- High-school maths: algebra with linear equations, the idea of a function \(f(x)\), slope of a straight line. You do not need prior multivariable calculus.
- Python: functions, lists, basic NumPy arrays (
np.array,@ornp.dot, broadcasting at a reading level). - Mindset: willingness to slow down at equations. Skipping the symbol explanations is how people stay stuck.
- Helpful but optional: skim section 1–2 of my AI & Machine Learning Foundations guide so "model, loss, train" already mean something.
Setup#
python -m venv .venv-maths
source .venv-maths/bin/activate # Windows: .venv-maths\Scripts\activate
pip install numpy
Versions used for every output in this tutorial:
| Package | Version |
|---|---|
| Python | 3.13.5 |
| NumPy | 2.5.3 |
No GPU. No API keys. Full lab script: maths_for_ml_lab.py — download the lab ZIP (script + README), or open it in the Hands-on lab section below.
Why ML needs maths (a concrete prediction)#
Suppose you want to predict an exam score from two numbers: hours studied and hours slept last night.
A very small model says:
| Symbol | Meaning |
|---|---|
| \(x_1, x_2\) | Feature values for one student (hours studied, hours slept) |
| \(w_1, w_2\) | Weights the model learned (how much each feature matters) |
| \(b\) | Bias (baseline score when features are zero — often not literally meaningful, but mathematically useful) |
| \(\hat{y}\) | Predicted score ("y-hat") |
EXAMPLE. Take \(x_1=4\), \(x_2=7\), \(w_1=8\), \(w_2=3\), \(b=20\):
That is already linear algebra: \((w_1, w_2)\) is a vector, \((x_1, x_2)\) is a vector, and \(w_1 x_1 + w_2 x_2\) is their dot product. Training the model means adjusting \(w_1, w_2, b\) so predictions get closer to real scores. Measuring "closer" needs a cost function. Improving the weights needs derivatives (or their vector form, the gradient). Taking a step downhill on that cost is gradient descent.
If you only remember one sentence from this page: ML maths is mostly "combine features with weights, measure error, move weights against the gradient of that error."
Core concept: three objects you will meet constantly#
| Object | What it is in ML | Mental picture |
|---|---|---|
| Vector | A list of numbers: a feature row, a weight vector, an embedding | An arrow in \(n\)-dimensional space |
| Matrix | A table of numbers: a batch of feature rows, a layer's weights | A stack of vectors / a linear map |
| Scalar | A single number: a loss value, a learning rate, one prediction | A point on the number line |
Beginners often confuse:
| Confused pair | Short fix |
|---|---|
| Scalar vs vector | Scalar = one number. Vector = ordered list. Loss is usually a scalar; features are a vector. |
| Row vs column | In this tutorial, feature vectors for one example are often written as columns in equations, but NumPy 1D arrays do not care until you stack them into a matrix. Always check shapes with .shape. |
| Derivative vs gradient | Derivative: slope for a function of one variable. Gradient: the vector of all partial derivatives for a function of many variables. |
| Dot product vs element-wise multiply | Dot product → one number. * in NumPy on same-shaped arrays → another array. |
Linear algebra for ML: vectors, matrices, operations you will use#
Vectors#
A vector in \(\mathbb{R}^n\) is an ordered list of \(n\) real numbers. We write column vectors in display math:
| Symbol | Meaning |
|---|---|
| \(\mathbf{x}\) | Bold = vector (common convention) |
| \(x_i\) | \(i\)-th component |
| \(\mathbb{R}^n\) | Space of \(n\)-dimensional real vectors |
In NumPy:
import numpy as np
x = np.array([4.0, 7.0]) # shape (2,) — a 1D vector
print(x.shape) # (2,)
Dot product (the weighted sum)#
The dot product of two vectors of equal length is:
| Symbol | Meaning |
|---|---|
| \(\mathbf{w} \cdot \mathbf{x}\) | Dot product (also written \(\mathbf{w}^\top \mathbf{x}\)) |
| \(\sum_{i=1}^{n}\) | Sum from \(i=1\) to \(n\) |
| \(w_i x_i\) | Product of corresponding components |
Why ML cares: a linear layer's forward pass on one example is a dot product (plus bias). Retrieval in RAG scores a query embedding against document embeddings with the same idea (often as cosine similarity, which is a normalised dot product).
Matrix-vector multiply as a batch of weighted sums#
Stack three students' features as rows of a matrix \(\mathbf{X}\):
Then:
Add bias \(b=20\) to each entry and you get predictions \([73, 83, 60]\) — the same numbers our lab prints below. Matrix-vector multiply = apply the same weighted sum to every row.
Code example — prediction as w·x + b#
import numpy as np
x = np.array([4.0, 7.0])
w = np.array([8.0, 3.0])
b = 20.0
y_hat = float(np.dot(w, x) + b)
print(y_hat)
X = np.array([[4.0, 7.0], [6.0, 5.0], [2.0, 8.0]])
print(X @ w + b)
Expected output (run 25 Sep 2026 IST, NumPy 2.5.3):
73.0
[73. 83. 60.]
Norms and distance (why L2 shows up)#
The L2 norm (Euclidean length) of a vector is:
| Symbol | Meaning |
|---|---|
| \(\|\mathbf{a}\|_2\) | L2 norm of \(\mathbf{a}\) |
| \(\sqrt{\cdot}\) | Square root |
For \(\mathbf{a} = (3, 4)\): \(\|\mathbf{a}\|_2 = \sqrt{9+16} = 5\).
Euclidean distance between \(\mathbf{a}\) and \(\mathbf{b}\) is the L2 norm of their difference:
The L1 norm (Manhattan length) is \(\|\mathbf{a}\|_1 = \sum_i |a_i|\). For \((3,-4)\) that is \(7\). Regularisation terms in ML (L1 / L2 penalties) are named after these norms.
Cosine similarity (preview for embeddings / RAG):
If \(\mathbf{u}\) and \(\mathbf{v}\) point the same way, cosine is \(1\). Parallel vectors \([1,2,3]\) and \([2,4,6]\) give exactly that in the lab.
a = np.array([3.0, 4.0])
print(np.linalg.norm(a, ord=2)) # 5.0
u = np.array([1.0, 2.0, 3.0])
v = np.array([2.0, 4.0, 6.0])
cos = np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v))
print(f"{cos:.4f}") # 1.0000
Calculus for ML: derivatives, partials, and gradients#
Derivative (one variable)#
If \(J(w)\) is a cost depending on one weight \(w\), the derivative \(\frac{dJ}{dw}\) is the slope: how much \(J\) changes if you nudge \(w\).
EXAMPLE. Let \(J(w) = (w - 3)^2\). Expand: \(J(w) = w^2 - 6w + 9\). Differentiate term by term:
| Symbol | Meaning |
|---|---|
| \(J(w)\) | Cost (also called loss) as a function of \(w\) |
| \(\frac{dJ}{dw}\) | Derivative of \(J\) with respect to \(w\) |
| \(w^*\) | Value of \(w\) that minimises \(J\) (here \(w^*=3\)) |
At \(w=3\) the slope is \(0\) — a minimum. For \(w<3\) the slope is negative; for \(w>3\) it is positive.
Partial derivatives (many variables)#
If \(J(w_1, w_2)\) depends on two weights, treat one as constant while differentiating the other:
| Symbol | Meaning |
|---|---|
| \(\partial\) | Partial derivative (only one variable moves) |
Gradient (the vector of all partials)#
| Symbol | Meaning |
|---|---|
| \(\nabla J\) | Gradient of \(J\) ("nabla J") |
| \(\mathbf{w}\) | Parameter vector \((w_1, w_2)\) |
Key fact used everywhere in ML: the gradient points in the direction of steepest ascent of \(J\). To decrease the cost, move the opposite way:
| Symbol | Meaning |
|---|---|
| \(\eta\) | Learning rate (step size; a positive scalar, e.g. \(0.1\)) |
| \(-\) | Minus sign: walk downhill, not uphill |
That update rule is gradient descent.
How it works: gradient descent on a simple cost#
Visual explanation#
- Start at a random \(w\)
- Compute the cost \(J(w)\)
- Compute the gradient \(\nabla J\)
- Update \(w \leftarrow w - \eta\,\nabla J\)
- Cost small enough? No: back to step 2. Yes: stop, \(w \approx w^*\)
Think of \(J\) as height on a landscape. You are standing at \(\mathbf{w}\). The gradient tells you which way the hill rises fastest. You face the opposite direction and take a step of length controlled by \(\eta\).
| Learning rate \(\eta\) | Typical behaviour |
|---|---|
| Too small | Crawl toward the minimum; many steps |
| Just right | Steady decrease in \(J\) |
| Too large | Overshoot; cost can increase and diverge |
Mathematical foundation — 1D walk-through#
Cost: \(J(w)=(w-3)^2\). Gradient (here just the derivative): \(2(w-3)\). Start at \(w=-2\), \(\eta=0.1\).
| Step | \(w\) | \(J(w)\) | \(dJ/dw\) |
|---|---|---|---|
| 0 | −2.000 | 25.0000 | −10.000 |
| 5 | 1.362 | 2.6844 | −3.277 |
| 10 | 2.463 | 0.2882 | −1.074 |
| 20 | 2.942 | 0.0033 | −0.115 |
| final | ≈2.981 | ≈0.00036 | ≈0 |
After 25 steps we are near \(w^*=3\). Same run as Part 3 of the lab.
Code example — 1D gradient descent#
def J(w: float) -> float:
return (w - 3.0) ** 2
def dJ(w: float) -> float:
return 2.0 * (w - 3.0)
w = -2.0
lr = 0.1
for step in range(25):
w = w - lr * dJ(w)
print(f"w = {w:.6f}, J = {J(w):.8f}")
Expected output:
w = 2.981111, J = 0.00035681
2D quadratic (anisotropic bowl)#
The minimum is at \(\mathbf{w}^*=(2,-1)\). The factor \(4\) makes the bowl steeper in the \(w_2\) direction — a tiny taste of why adaptive optimisers exist later (not covered here).
With \(\mathbf{w}_0=(0,0)\), \(\eta=0.05\), after 40 steps the lab reaches approximately \([1.970, -1.000]\) with \(J \approx 0.00087\).
Practical example: linear regression training is the same idea#
For \(n\) examples with features in matrix \(\mathbf{X}\in\mathbb{R}^{n\times d}\) (we add a column of ones so the bias sits inside \(\boldsymbol{\theta}\)) and targets \(\mathbf{y}\in\mathbb{R}^n\):
Mean squared error (MSE):
| Symbol | Meaning |
|---|---|
| \(\boldsymbol{\theta}\) | Parameter vector (weights + bias) |
| \(\hat{y}_i\) | Prediction for example \(i\) |
| \(y_i\) | True target for example \(i\) |
Gradient (Level 3 form you will see in notes):
| Symbol | Meaning |
|---|---|
| \(\mathbf{X}^\top\) | Transpose of \(\mathbf{X}\) (flips rows/columns) |
| \((\mathbf{X}\boldsymbol{\theta} - \mathbf{y})\) | Vector of residuals (prediction − truth) |
Update: \(\boldsymbol{\theta} \leftarrow \boldsymbol{\theta} - \eta \nabla_{\boldsymbol{\theta}} J\).
EXAMPLE (synthetic). True rule \(y = 5x_1 + 2x_2 + 1\) plus small noise. After 80 steps with \(\eta=0.1\), NumPy recovered \(\boldsymbol{\theta} \approx [4.990, 2.001, 0.931]\) — close to \([5, 2, 1]\). Residual MSE ≈ \(0.054\) (noise floor; not zero).
Preview: neural net training#
A neural net stacks many such linear maps with non-linear activations between them. The cost is still a scalar. Backpropagation is the chain rule applied to compute \(\nabla_{\boldsymbol{\theta}} J\) for every parameter. The update rule is still "subtract learning rate times gradient." If you understand this page, the first half of a deep-learning chapter stops looking like magic.
We are not implementing a net here. We are making sure the optimisation story is solid before frameworks hide it.
How It Works (end-to-end story in one place)#
Here is the same pipeline without new notation, so you can rehearse it aloud before an interview.
- Represent one example as a feature vector \(\mathbf{x}\). Represent the model as weights \(\mathbf{w}\) and bias \(b\) (or one augmented \(\boldsymbol{\theta}\)).
- Predict with a weighted sum: \(\hat{y} = \mathbf{w}\cdot\mathbf{x} + b\). For a batch, use \(\mathbf{X}\mathbf{w}\).
- Score the error with a cost such as MSE. The cost is a scalar — one number that summarises how unhappy you are with the current weights.
- Differentiate that scalar with respect to every parameter. Bundle those partial derivatives into \(\nabla J\).
- Step opposite the gradient: subtract \(\eta\nabla J\). Repeat until the cost flattens or you hit a step budget.
- Read the result. If features were meaningful and the model class was rich enough, \(\mathbf{w}\) often has an interpretation (e.g. "each extra study hour adds about 5 points"). If you used a deep net, individual weights are harder to read, but the training loop is still steps 3–5.
That loop is why calculus shows up in ML even when your day job is mostly Python.
Level ladder used in this tutorial#
| Level | What you saw | Example |
|---|---|---|
| 1 — Intuition | Prediction = weighted sum; training = walk downhill on error | Exam-score story |
| 2 — Definitions | Vector, matrix, norm, derivative, gradient | Symbol tables after each equation |
| 3 — Mathematical form | MSE gradient \(\frac{2}{n}\mathbf{X}^\top(\mathbf{X}\boldsymbol{\theta}-\mathbf{y})\) | Linear regression section |
| 4 — From-scratch code | NumPy loops, no scikit-learn | maths_for_ml_lab.py Parts 3–5 |
We deliberately stop before Level-5 "call a library estimator," so the idea is not skipped.
Worked comparison: three learning rates on the noisy regression#
Same Part-5 data, 80 steps, three values of \(\eta\):
| \(\eta\) | Final MSE (measured) | What you should notice |
|---|---|---|
| 0.001 | 17.93 | Barely moved; weights still near zero |
| 0.1 | 0.054 | Settled near the noise floor; recovered weights ≈ true rule |
| 1.0 | 24.44 | Unstable / poor fit in this short run — not "faster learning" |
ASSUMPTION: these numbers are for this synthetic draw (numpy seed 42), not a universal law. The qualitative lesson is stable: tune \(\eta\) by watching the cost curve, not by folklore.
Shape cheat for the regression gradient#
For \(\mathbf{X}\) shaped \((n, d)\), \(\boldsymbol{\theta}\) shaped \((d,)\), \(\mathbf{y}\) shaped \((n,)\):
| Expression | Shape | Role |
|---|---|---|
| \(\mathbf{X}\boldsymbol{\theta}\) | \((n,)\) | Predictions |
| \(\mathbf{X}\boldsymbol{\theta}-\mathbf{y}\) | \((n,)\) | Residuals |
| \(\mathbf{X}^\top(\mathbf{X}\boldsymbol{\theta}-\mathbf{y})\) | \((d,)\) | Unnormalised gradient direction |
| \(\frac{2}{n}\mathbf{X}^\top(\cdots)\) | \((d,)\) | MSE gradient |
If a line in your notebook throws matmul shape errors, fix shapes before doubting the formula.
Hands-on lab (45–60 minutes)#
Goal: run and extend maths_for_ml_lab.py so every abstract symbol has a number next to it.
- Create the venv (see Setup) and run:
python maths_for_ml_lab.py
- Part 1 — confirm prediction
73.0and batch[73. 83. 60.]by hand. - Part 2 — confirm L2 of \((3,4)\) is \(5\); cosine of parallel vectors is \(1\).
- Part 3 — watch \(w\) climb from \(-2\) toward \(3\). Change
lrto1.1and observe divergence (see Common Mistakes). - Part 4 — note that \(w_2\) reaches \(-1\) faster than \(w_1\) reaches \(2\) (steeper valley).
- Part 5 — recover weights near \([5, 2, 1]\). Try
lr=0.001andlr=1.0; record final MSE.
Lab outputs (abridged, real run):
Prediction y_hat = w·x + b = 73.0
Batch predictions X @ w + b = [73. 83. 60.]
||a||_2 = 5.0
cosine_similarity = 1.0000
final 1D: w = 2.981111, J(w) = 0.00035681
final 2D: w = [1.970438, -1.000000], J = 0.00087390
Recovered theta ≈ [4.990, 2.001, 0.931]
View the full lab script (maths_for_ml_lab.py)
#!/usr/bin/env python3
"""
Mathematics for Machine Learning — Hands-on Lab
Companion to: Mathematics for Machine Learning: The Essentials You'll Actually Use
Site: pushpjeet.com/tutorials/mathematics-for-machine-learning/
Run: python maths_for_ml_lab.py
Needs: Python 3.10+, NumPy. No GPU, no API keys.
Estimated time: 45–60 minutes if you type the exercises yourself.
"""
from __future__ import annotations
import numpy as np
np.set_printoptions(precision=4, suppress=True)
def section(title: str) -> None:
print("\n" + "=" * 60)
print(title)
print("=" * 60)
# ---------------------------------------------------------------------------
# Part 1 — Vectors, matrices, dot product (prediction as weighted sum)
# ---------------------------------------------------------------------------
def part1_vectors_and_dot_product() -> None:
section("Part 1: Vectors, matrices, and a tiny prediction")
# EXAMPLE: hours studied (x1) and hours slept (x2) -> predicted exam score
# Model: y_hat = w1*x1 + w2*x2 + b
x = np.array([4.0, 7.0]) # features for one student
w = np.array([8.0, 3.0]) # learned weights
b = 20.0 # bias
y_hat = float(np.dot(w, x) + b)
print(f"Feature vector x = {x}")
print(f"Weight vector w = {w}")
print(f"Bias b = {b}")
print(f"Prediction y_hat = w·x + b = {y_hat}")
print("(Hand check: 8*4 + 3*7 + 20 = 32 + 21 + 20 = 73)")
# Batch of 3 students as rows of a matrix X (shape 3 x 2)
X = np.array([
[4.0, 7.0],
[6.0, 5.0],
[2.0, 8.0],
])
y_batch = X @ w + b # matrix-vector multiply = weighted sums
print(f"\nBatch feature matrix X (3 students x 2 features):\n{X}")
print(f"Batch predictions X @ w + b = {y_batch}")
# ---------------------------------------------------------------------------
# Part 2 — Norms and distance
# ---------------------------------------------------------------------------
def part2_norms_and_distance() -> None:
section("Part 2: L2 norm and Euclidean distance")
a = np.array([3.0, 4.0])
b = np.array([0.0, 0.0])
l2_a = float(np.linalg.norm(a, ord=2))
dist = float(np.linalg.norm(a - b, ord=2))
print(f"a = {a}")
print(f"||a||_2 = sqrt(3^2 + 4^2) = {l2_a}")
print(f"distance(a, origin) = {dist}")
# Cosine similarity (used later in RAG / embeddings)
u = np.array([1.0, 2.0, 3.0])
v = np.array([2.0, 4.0, 6.0]) # parallel to u
cos = float(np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v)))
print(f"\ncosine_similarity(u, v) for parallel vectors = {cos:.4f} (expect 1.0)")
# ---------------------------------------------------------------------------
# Part 3 — Gradient of a simple quadratic cost (1D)
# ---------------------------------------------------------------------------
def part3_1d_gradient_descent() -> None:
section("Part 3: Gradient descent on J(w) = (w - 3)^2 [1D]")
# J(w) = (w - 3)^2
# dJ/dw = 2(w - 3)
# Minimum at w* = 3, J* = 0
def J(w: float) -> float:
return (w - 3.0) ** 2
def dJ(w: float) -> float:
return 2.0 * (w - 3.0)
w = -2.0
lr = 0.1
history = []
for step in range(25):
history.append((step, w, J(w), dJ(w)))
w = w - lr * dJ(w)
print("step | w | J(w) | dJ/dw")
print("-----+--------+----------+-------")
for step, ww, jj, g in history[::5]: # every 5 steps
print(f"{step:4d} | {ww:6.3f} | {jj:8.4f} | {g:6.3f}")
print(f"final: w = {w:.6f}, J(w) = {J(w):.8f} (target w*=3)")
# ---------------------------------------------------------------------------
# Part 4 — Gradient descent on a 2D quadratic cost
# ---------------------------------------------------------------------------
def part4_2d_gradient_descent() -> dict:
section("Part 4: Gradient descent on J(w) = (w1-2)^2 + 4*(w2+1)^2 [2D]")
# Minimum at w* = (2, -1)
# grad J = [2(w1-2), 8(w2+1)]
def J(w: np.ndarray) -> float:
return float((w[0] - 2.0) ** 2 + 4.0 * (w[1] + 1.0) ** 2)
def grad_J(w: np.ndarray) -> np.ndarray:
return np.array([2.0 * (w[0] - 2.0), 8.0 * (w[1] + 1.0)])
w = np.array([0.0, 0.0])
lr = 0.05
path = [w.copy()]
costs = [J(w)]
for step in range(40):
g = grad_J(w)
w = w - lr * g
path.append(w.copy())
costs.append(J(w))
print("step | w1 | w2 | J(w)")
print("-----+--------+--------+--------")
for i in [0, 5, 10, 20, 39]:
ww = path[i]
print(f"{i:4d} | {ww[0]:6.3f} | {ww[1]:6.3f} | {costs[i]:7.4f}")
print(f"final: w = [{w[0]:.6f}, {w[1]:.6f}], J = {J(w):.8f}")
print("target: w* = [2.0, -1.0], J* = 0")
return {"final_w": w, "final_J": J(w), "path": path, "costs": costs}
# ---------------------------------------------------------------------------
# Part 5 — Tiny linear regression from scratch (2 features + bias)
# ---------------------------------------------------------------------------
def part5_linear_regression_gd() -> None:
section("Part 5: Linear regression with gradient descent (from scratch)")
# Synthetic data: y = 5*x1 + 2*x2 + 1 + small noise
rng = np.random.default_rng(42)
n = 40
X_raw = rng.normal(size=(n, 2))
true_w = np.array([5.0, 2.0])
true_b = 1.0
y = X_raw @ true_w + true_b + rng.normal(scale=0.3, size=n)
# Augment with a column of ones for the bias (design matrix)
X = np.column_stack([X_raw, np.ones(n)]) # shape (n, 3)
# Parameters: theta = [w1, w2, b]
theta = np.zeros(3)
lr = 0.1
history = []
for step in range(80):
preds = X @ theta
err = preds - y
# Mean squared error
mse = float(np.mean(err ** 2))
# Gradient of MSE w.r.t. theta: (2/n) X^T (X theta - y)
grad = (2.0 / n) * (X.T @ err)
theta = theta - lr * grad
if step % 20 == 0 or step == 79:
history.append((step, theta.copy(), mse))
print("True weights: w=[5, 2], b=1")
print("step | w1 | w2 | b | MSE")
print("-----+--------+--------+--------+--------")
for step, th, mse in history:
print(f"{step:4d} | {th[0]:6.3f} | {th[1]:6.3f} | {th[2]:6.3f} | {mse:7.4f}")
print(f"\nRecovered theta ≈ [{theta[0]:.3f}, {theta[1]:.3f}, {theta[2]:.3f}]")
# ---------------------------------------------------------------------------
# Student exercises (stubs — solutions printed when --solutions is passed)
# ---------------------------------------------------------------------------
def exercise_hints() -> None:
section("Exercises (do these yourself; answers at end of tutorial)")
print(
"""
E1. Change the learning rate in Part 3 to 1.1. What happens? Why?
E2. Implement L1 norm of a = [3, -4] by hand and with np.linalg.norm(a, ord=1).
E3. For J(w1,w2) = w1^2 + w2^2, write the gradient by hand, then verify with
a tiny finite-difference check in NumPy (perturb each coordinate by 1e-5).
E4. In Part 5, try lr=0.001 and lr=1.0. Record final MSE after 80 steps for each.
E5. Explain in one sentence: why do we subtract the gradient (not add it)?
"""
)
def main() -> None:
print("Mathematics for Machine Learning — Lab")
print("Python / NumPy from-scratch examples (no scikit-learn)")
part1_vectors_and_dot_product()
part2_norms_and_distance()
part3_1d_gradient_descent()
part4_2d_gradient_descent()
part5_linear_regression_gd()
exercise_hints()
print("\nLab finished. Paste key outputs into your notes.")
if __name__ == "__main__":
main()
Download the lab ZIPCheat sheet PDFCheat sheet (web)
A one-page symbol sheet lives in the companion student cheat sheet (also as a printable PDF).
Common mistakes#
-
Learning rate too large. On \(J(w)=(w-3)^2\) with \(w_0=-2\) and \(\eta=1.1\), cost grows: \(25 \rightarrow 36 \rightarrow 52 \rightarrow \ldots \rightarrow 321\) by step 7. The step jumps over the minimum and climbs the other side. Fix: reduce \(\eta\) (try \(0.1\), \(0.01\)).
-
Adding the gradient instead of subtracting. That maximises the cost. Training "works" only if you minimise — always remember the minus sign.
-
Shape / transpose bugs.
X @ wneeds matching inner dimensions. Print.shapebefore debugging maths that is actually fine. -
Forgetting the bias. A model forced through the origin cannot fit \(y = 5x + 1\). Augment \(\mathbf{X}\) with a column of ones or keep \(b\) separate and update it with \(\frac{\partial J}{\partial b}\).
-
Confusing L1 and L2. L2 squares then square-roots; L1 sums absolute values. Regularisation and "distance" mean different geometries.
-
Treating training loss \(0\) as always possible. With noisy labels (Part 5), MSE bottoms near the noise level (~\(0.05\) here), not zero. Chasing zero then overfits.
-
Jumping to
sklearn.LinearRegressionbefore deriving the gradient once. Libraries are fine in production; they are a poor first teacher. Derive, implement once, then use the library.
Real-world applications#
| Area | Where this maths appears |
|---|---|
| Tabular ML | Linear / logistic regression, feature scaling (often L2-related), regularisation |
| Embeddings & RAG | Dot products and cosine similarity between vectors — see the RAG tutorial |
| Neural nets | Every dense layer is matrix multiply; training is gradient descent (SGD / Adam variants) |
| Recommendation | User and item vectors; scores as dots |
| Computer vision / NLP | Same linear algebra at huge scale inside convolution and attention (attention scores are scaled dots) |
Interview questions#
Q1. What is the difference between a derivative and a gradient?
A: A derivative is the slope of a function of one variable. A gradient is the vector of partial derivatives of a function of several variables. Each component says how the function changes if you nudge that one parameter.
Q2. Why do we subtract the gradient in gradient descent?
A: The gradient points to steepest ascent of the cost. Subtracting it moves parameters toward lower cost. The learning rate scales the step.
Q3. Express a linear prediction with a bias using a dot product.
A: Either \(\hat{y}=\mathbf{w}\cdot\mathbf{x}+b\), or fold \(b\) into an augmented vector: \(\hat{y}=\boldsymbol{\theta}\cdot[x_1,\ldots,x_n,1]\).
Q4. What happens if the learning rate is too large?
A: Updates overshoot the minimum; the loss can oscillate or diverge. (Quote the \(\eta=1.1\) numbers from Common Mistakes if whiteboarding.)
Q5. How is cosine similarity related to the dot product?
A: Cosine is the dot product divided by the product of L2 norms. On unit vectors, cosine equals the dot product — which is why many retrieval stacks normalise embeddings first (RAG tutorial).
MCQs (10)#
1. The expression \(w_1 x_1 + w_2 x_2\) is best described as:
- Element-wise product
- Dot product
- Matrix inverse
- Determinant
Show answer
Answer: B
2. In \(\mathbf{w} \leftarrow \mathbf{w} - \eta \nabla J\), \(\eta\) is:
- Momentum
- Learning rate
- Regularisation strength
- Batch size
Show answer
Answer: B
3. \(\|\mathbf{a}\|_2\) for \(\mathbf{a}=(3,4)\) equals:
- 7
- 12
- 5
- 25
Show answer
Answer: C
4. The gradient \(\nabla J\) points toward:
- Steepest descent
- Steepest ascent
- A random direction
- Always the origin
Show answer
Answer: B
5. MSE for regression is:
- Mean of absolute residuals
- Mean of squared residuals
- Max residual
- Median residual
Show answer
Answer: B
6. For \(J(w)=(w-3)^2\), the derivative \(dJ/dw\) is:
- \(w-3\)
- \(2(w-3)\)
- \(2w\)
- \(3-w\)
Show answer
Answer: B
7. Adding a column of ones to \(\mathbf{X}\) before \(\mathbf{X}\boldsymbol{\theta}\) mainly helps you:
- Normalise features
- Encode the bias inside \(\boldsymbol{\theta}\)
- Compute determinants
- Shuffle rows
Show answer
Answer: B
8. Cosine similarity of two parallel non-zero vectors is:
- 0
- −1
- 1
- Undefined
Show answer
Answer: C
9. L1 norm of \((3,-4)\) is:
- 5
- 7
- 12
- 1
Show answer
Answer: B
10. If training MSE stops near \(0.05\) on noisy data, the most likely explanation is:
- Bug in NumPy
- Noise floor / irreducible error
- Learning rate must be 1.0
- Need a deeper net always
Show answer
Answer: B
Practice exercises#
E1. In Part 3 of the lab, set lr = 1.1. Record \(w\) and \(J\) for 8 steps. Explain the growth using the update rule.
(Measured: \(J\) goes \(25 \rightarrow 36 \rightarrow 52 \rightarrow 75 \rightarrow 107 \rightarrow 155 \rightarrow 223 \rightarrow 321\).)
E2. Compute \(\|\mathbf{a}\|_1\) for \(\mathbf{a}=(3,-4)\) by hand and with np.linalg.norm(a, ord=1).
(Both give \(7.0\).)
E3. For \(J(w_1,w_2)=w_1^2+w_2^2\), write \(\nabla J\). At \((1.5,-0.8)\) check with central finite differences (\(\varepsilon=10^{-5}\)).
(Analytic and finite-difference both \([3.0,\,-1.6]\).)
E4. In Part 5, after 80 steps record final MSE for \(\eta\in\{0.001,\,0.1,\,1.0\}\).
(Measured: \(17.93\), \(0.054\), \(24.44\) respectively — too small crawls; too large fails to settle near the optimum in this setup.)
E5. In one sentence: why subtract the gradient?
(Because the gradient points uphill on the cost; we want to go downhill.)
Downloads#
Free companion material for this tutorial. The numbers in all three match the lab run shown on this page (Python 3.13.5, NumPy 2.5.3).
Student cheat sheet (PDF)
Symbols, must-know equations, NumPy patterns and lab checkpoints on one printable sheet.
Download PDFHands-on lab (ZIP)
maths_for_ml_lab.py plus a short README: vectors, norms, 1D/2D gradient descent and linear regression from scratch.
Summary#
- A linear prediction is a dot product of weights and features, plus bias.
- Matrices apply that weighted sum to a whole batch at once.
- L2 norms measure length and distance; cosine is a normalised dot product (RAG retrieval uses this).
- Derivatives give slope in 1D; gradients package all partial derivatives in higher dimensions.
- Gradient descent repeatedly walks opposite the gradient, scaled by the learning rate.
- Linear regression training and (at a high level) neural net training share that optimisation story.
- Implement the idea once in NumPy. Then frameworks make sense instead of feeling like superstition.
Further learning (primary / well-known sources)#
- 3Blue1Brown — Essence of linear algebra (visual intuition for vectors, dot products, matrices). Conceptual further learning; watch before heavier texts.
- Gilbert Strang — Introduction to Linear Algebra (and MIT OCW linear algebra lectures). Classic, clear on the computations ML reuses.
- Ian Goodfellow, Yoshua Bengio, Aaron Courville — Deep Learning (MIT Press): early chapters on linear algebra, probability, and numerical computation for deep learning.
- Mathematics for Machine Learning (Deisenroth, Faisal, Ong) — free textbook aimed exactly at this bridge; use as a slower parallel track after this pillar.
- Pushpjeet — AI & Machine Learning Foundations and RAG Tutorial in Python for where this maths shows up in systems.
FAQ#
Is this enough maths to start deep learning?
Enough to understand loss, gradients, and why training moves weights. You will still need the chain rule / backprop in more detail for deep nets — this page is the on-ramp, not the whole highway.
Do I need to memorise matrix calculus identities?
Memorise the few you use weekly (dot product, MSE gradient for linear regression, gradient descent update). Look up the rest. Understanding beats memorising forty identities.
Why NumPy and not scikit-learn first?
Because LinearRegression().fit hides the gradient. Once you have computed \(\nabla J\) yourself, using scikit-learn is a speed choice, not a conceptual leap.
How does this connect to RAG?
Document and query embeddings are vectors. Ranking by cosine similarity is ranking by a normalised dot product — the same linear algebra as Part 1–2, applied to text. Details: /tutorials/rag-tutorial-python/.
Does this page cover probability, Bayes or softmax?
No. This page stays on the linear algebra and calculus essentials for training. Probability is a separate topic; join the newsletter if you want to hear when new tutorials are published.
Continue
Practise it: the cheat sheet (or the PDF) and the lab ZIP use the same verified numbers.
Related: Linear Regression for Beginners · Linear Regression (Expert) · AI & ML Foundations · RAG Tutorial in Python
Learn live: Training · WhatsApp +91 70492 35525