← pushpjeet.com · TutorialsCheat sheet PDF
Intermediate beginnerLinear algebraGradient descentNumPy lab

Mathematics for Machine Learning: The Essentials You'll Actually Use

You can call model.fit without understanding a single equation. Then an interview asks why the loss blew up when the learning rate changed, or why cosine similarity prefers unit vectors, and the gap shows.

This tutorial is the bridge: the mathematics for machine learning you will actually use — vectors and matrices as features and weights, norms as distances, derivatives and gradients as "which way reduces error", and gradient descent as the workhorse that trains almost everything from linear regression to neural nets. We stay practical. Every important equation is written in LaTeX, every symbol is named right after it, and every key idea is followed by NumPy you can run.

We will not cover measure theory, proofs of the spectral theorem, or a full course in multivariable calculus. We will cover what you need to read a training loop without guessing.

Learning objectives#

By the end you will be able to:

  1. Explain why a prediction is a weighted sum of features (dot product / matrix-vector multiply).
  2. Distinguish scalar vs vector, row vs column, derivative vs gradient, and L1 vs L2 norms in plain language.
  3. Compute a dot product, an L2 norm, and a Euclidean distance by hand and in NumPy.
  4. Write the gradient of a simple cost and explain why it points toward steepest ascent.
  5. Implement gradient descent from scratch on a 1D and a 2D quadratic cost, and on a tiny linear regression problem.
  6. Connect that maths to linear regression training and to the first layer of a neural net (preview only).
  7. Spot common mistakes: learning rate too large, confusing transpose shapes, forgetting the bias term.

Prerequisites#

  • High-school maths: algebra with linear equations, the idea of a function \(f(x)\), slope of a straight line. You do not need prior multivariable calculus.
  • Python: functions, lists, basic NumPy arrays (np.array, @ or np.dot, broadcasting at a reading level).
  • Mindset: willingness to slow down at equations. Skipping the symbol explanations is how people stay stuck.
  • Helpful but optional: skim section 1–2 of my AI & Machine Learning Foundations guide so "model, loss, train" already mean something.

Setup#

python -m venv .venv-maths
source .venv-maths/bin/activate   # Windows: .venv-maths\Scripts\activate
pip install numpy

Versions used for every output in this tutorial:

Package Version
Python 3.13.5
NumPy 2.5.3

No GPU. No API keys. Full lab script: maths_for_ml_lab.py — download the lab ZIP (script + README), or open it in the Hands-on lab section below.

Why ML needs maths (a concrete prediction)#

Suppose you want to predict an exam score from two numbers: hours studied and hours slept last night.

A very small model says:

\[\hat{y} = w_1 x_1 + w_2 x_2 + b\]
Symbol Meaning
\(x_1, x_2\) Feature values for one student (hours studied, hours slept)
\(w_1, w_2\) Weights the model learned (how much each feature matters)
\(b\) Bias (baseline score when features are zero — often not literally meaningful, but mathematically useful)
\(\hat{y}\) Predicted score ("y-hat")

EXAMPLE. Take \(x_1=4\), \(x_2=7\), \(w_1=8\), \(w_2=3\), \(b=20\):

\[\hat{y} = 8\cdot 4 + 3\cdot 7 + 20 = 32 + 21 + 20 = 73\]

That is already linear algebra: \((w_1, w_2)\) is a vector, \((x_1, x_2)\) is a vector, and \(w_1 x_1 + w_2 x_2\) is their dot product. Training the model means adjusting \(w_1, w_2, b\) so predictions get closer to real scores. Measuring "closer" needs a cost function. Improving the weights needs derivatives (or their vector form, the gradient). Taking a step downhill on that cost is gradient descent.

If you only remember one sentence from this page: ML maths is mostly "combine features with weights, measure error, move weights against the gradient of that error."

Core concept: three objects you will meet constantly#

Object What it is in ML Mental picture
Vector A list of numbers: a feature row, a weight vector, an embedding An arrow in \(n\)-dimensional space
Matrix A table of numbers: a batch of feature rows, a layer's weights A stack of vectors / a linear map
Scalar A single number: a loss value, a learning rate, one prediction A point on the number line

Beginners often confuse:

Confused pair Short fix
Scalar vs vector Scalar = one number. Vector = ordered list. Loss is usually a scalar; features are a vector.
Row vs column In this tutorial, feature vectors for one example are often written as columns in equations, but NumPy 1D arrays do not care until you stack them into a matrix. Always check shapes with .shape.
Derivative vs gradient Derivative: slope for a function of one variable. Gradient: the vector of all partial derivatives for a function of many variables.
Dot product vs element-wise multiply Dot product → one number. * in NumPy on same-shaped arrays → another array.

Linear algebra for ML: vectors, matrices, operations you will use#

Vectors#

A vector in \(\mathbb{R}^n\) is an ordered list of \(n\) real numbers. We write column vectors in display math:

\[\mathbf{x} = \begin{bmatrix} x_1 \\ x_2 \\ \vdots \\ x_n \end{bmatrix}\]
Symbol Meaning
\(\mathbf{x}\) Bold = vector (common convention)
\(x_i\) \(i\)-th component
\(\mathbb{R}^n\) Space of \(n\)-dimensional real vectors

In NumPy:

import numpy as np
x = np.array([4.0, 7.0])   # shape (2,) — a 1D vector
print(x.shape)             # (2,)

Dot product (the weighted sum)#

The dot product of two vectors of equal length is:

\[\mathbf{w} \cdot \mathbf{x} = \sum_{i=1}^{n} w_i x_i = w_1 x_1 + w_2 x_2 + \cdots + w_n x_n\]
Symbol Meaning
\(\mathbf{w} \cdot \mathbf{x}\) Dot product (also written \(\mathbf{w}^\top \mathbf{x}\))
\(\sum_{i=1}^{n}\) Sum from \(i=1\) to \(n\)
\(w_i x_i\) Product of corresponding components

Why ML cares: a linear layer's forward pass on one example is a dot product (plus bias). Retrieval in RAG scores a query embedding against document embeddings with the same idea (often as cosine similarity, which is a normalised dot product).

Matrix-vector multiply as a batch of weighted sums#

Stack three students' features as rows of a matrix \(\mathbf{X}\):

\[\mathbf{X} = \begin{bmatrix} 4 & 7 \\ 6 & 5 \\ 2 & 8 \end{bmatrix}, \quad \mathbf{w} = \begin{bmatrix} 8 \\ 3 \end{bmatrix}\]

Then:

\[\mathbf{X}\mathbf{w} = \begin{bmatrix} 4\cdot8 + 7\cdot3 \\ 6\cdot8 + 5\cdot3 \\ 2\cdot8 + 8\cdot3 \end{bmatrix} = \begin{bmatrix} 53 \\ 63 \\ 40 \end{bmatrix}\]

Add bias \(b=20\) to each entry and you get predictions \([73, 83, 60]\) — the same numbers our lab prints below. Matrix-vector multiply = apply the same weighted sum to every row.

Code example — prediction as w·x + b#

import numpy as np

x = np.array([4.0, 7.0])
w = np.array([8.0, 3.0])
b = 20.0
y_hat = float(np.dot(w, x) + b)
print(y_hat)

X = np.array([[4.0, 7.0], [6.0, 5.0], [2.0, 8.0]])
print(X @ w + b)

Expected output (run 25 Sep 2026 IST, NumPy 2.5.3):

73.0
[73. 83. 60.]

Norms and distance (why L2 shows up)#

The L2 norm (Euclidean length) of a vector is:

\[\|\mathbf{a}\|_2 = \sqrt{a_1^2 + a_2^2 + \cdots + a_n^2}\]
Symbol Meaning
\(\|\mathbf{a}\|_2\) L2 norm of \(\mathbf{a}\)
\(\sqrt{\cdot}\) Square root

For \(\mathbf{a} = (3, 4)\): \(\|\mathbf{a}\|_2 = \sqrt{9+16} = 5\).

Euclidean distance between \(\mathbf{a}\) and \(\mathbf{b}\) is the L2 norm of their difference:

\[d(\mathbf{a}, \mathbf{b}) = \|\mathbf{a} - \mathbf{b}\|_2\]

The L1 norm (Manhattan length) is \(\|\mathbf{a}\|_1 = \sum_i |a_i|\). For \((3,-4)\) that is \(7\). Regularisation terms in ML (L1 / L2 penalties) are named after these norms.

Cosine similarity (preview for embeddings / RAG):

\[\cos(\mathbf{u}, \mathbf{v}) = \frac{\mathbf{u}\cdot\mathbf{v}}{\|\mathbf{u}\|_2 \|\mathbf{v}\|_2}\]

If \(\mathbf{u}\) and \(\mathbf{v}\) point the same way, cosine is \(1\). Parallel vectors \([1,2,3]\) and \([2,4,6]\) give exactly that in the lab.

a = np.array([3.0, 4.0])
print(np.linalg.norm(a, ord=2))  # 5.0
u = np.array([1.0, 2.0, 3.0])
v = np.array([2.0, 4.0, 6.0])
cos = np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v))
print(f"{cos:.4f}")  # 1.0000

Calculus for ML: derivatives, partials, and gradients#

Derivative (one variable)#

If \(J(w)\) is a cost depending on one weight \(w\), the derivative \(\frac{dJ}{dw}\) is the slope: how much \(J\) changes if you nudge \(w\).

EXAMPLE. Let \(J(w) = (w - 3)^2\). Expand: \(J(w) = w^2 - 6w + 9\). Differentiate term by term:

\[\frac{dJ}{dw} = 2w - 6 = 2(w - 3)\]
Symbol Meaning
\(J(w)\) Cost (also called loss) as a function of \(w\)
\(\frac{dJ}{dw}\) Derivative of \(J\) with respect to \(w\)
\(w^*\) Value of \(w\) that minimises \(J\) (here \(w^*=3\))

At \(w=3\) the slope is \(0\) — a minimum. For \(w<3\) the slope is negative; for \(w>3\) it is positive.

Partial derivatives (many variables)#

If \(J(w_1, w_2)\) depends on two weights, treat one as constant while differentiating the other:

\[J(w_1, w_2) = (w_1 - 2)^2 + 4(w_2 + 1)^2\]
\[\frac{\partial J}{\partial w_1} = 2(w_1 - 2), \qquad \frac{\partial J}{\partial w_2} = 8(w_2 + 1)\]
Symbol Meaning
\(\partial\) Partial derivative (only one variable moves)

Gradient (the vector of all partials)#

\[\nabla J(\mathbf{w}) = \begin{bmatrix} \dfrac{\partial J}{\partial w_1} \\ \dfrac{\partial J}{\partial w_2} \end{bmatrix}\]
Symbol Meaning
\(\nabla J\) Gradient of \(J\) ("nabla J")
\(\mathbf{w}\) Parameter vector \((w_1, w_2)\)

Key fact used everywhere in ML: the gradient points in the direction of steepest ascent of \(J\). To decrease the cost, move the opposite way:

\[\mathbf{w}_{\text{new}} = \mathbf{w}_{\text{old}} - \eta \, \nabla J(\mathbf{w}_{\text{old}})\]
Symbol Meaning
\(\eta\) Learning rate (step size; a positive scalar, e.g. \(0.1\))
\(-\) Minus sign: walk downhill, not uphill

That update rule is gradient descent.

How it works: gradient descent on a simple cost#

Visual explanation#

Think of \(J\) as height on a landscape. You are standing at \(\mathbf{w}\). The gradient tells you which way the hill rises fastest. You face the opposite direction and take a step of length controlled by \(\eta\).

Learning rate \(\eta\) Typical behaviour
Too small Crawl toward the minimum; many steps
Just right Steady decrease in \(J\)
Too large Overshoot; cost can increase and diverge

Mathematical foundation — 1D walk-through#

Cost: \(J(w)=(w-3)^2\). Gradient (here just the derivative): \(2(w-3)\). Start at \(w=-2\), \(\eta=0.1\).

Step \(w\) \(J(w)\) \(dJ/dw\)
0 −2.000 25.0000 −10.000
5 1.362 2.6844 −3.277
10 2.463 0.2882 −1.074
20 2.942 0.0033 −0.115
final ≈2.981 ≈0.00036 ≈0

After 25 steps we are near \(w^*=3\). Same run as Part 3 of the lab.

Code example — 1D gradient descent#

def J(w: float) -> float:
    return (w - 3.0) ** 2

def dJ(w: float) -> float:
    return 2.0 * (w - 3.0)

w = -2.0
lr = 0.1
for step in range(25):
    w = w - lr * dJ(w)
print(f"w = {w:.6f}, J = {J(w):.8f}")

Expected output:

w = 2.981111, J = 0.00035681

2D quadratic (anisotropic bowl)#

\[J(w_1, w_2) = (w_1 - 2)^2 + 4(w_2 + 1)^2, \quad \nabla J = \begin{bmatrix} 2(w_1-2) \\ 8(w_2+1) \end{bmatrix}\]

The minimum is at \(\mathbf{w}^*=(2,-1)\). The factor \(4\) makes the bowl steeper in the \(w_2\) direction — a tiny taste of why adaptive optimisers exist later (not covered here).

With \(\mathbf{w}_0=(0,0)\), \(\eta=0.05\), after 40 steps the lab reaches approximately \([1.970, -1.000]\) with \(J \approx 0.00087\).

Practical example: linear regression training is the same idea#

For \(n\) examples with features in matrix \(\mathbf{X}\in\mathbb{R}^{n\times d}\) (we add a column of ones so the bias sits inside \(\boldsymbol{\theta}\)) and targets \(\mathbf{y}\in\mathbb{R}^n\):

\[\hat{\mathbf{y}} = \mathbf{X}\boldsymbol{\theta}\]

Mean squared error (MSE):

\[J(\boldsymbol{\theta}) = \frac{1}{n}\sum_{i=1}^{n}(\hat{y}_i - y_i)^2 = \frac{1}{n}\|\mathbf{X}\boldsymbol{\theta} - \mathbf{y}\|_2^2\]
Symbol Meaning
\(\boldsymbol{\theta}\) Parameter vector (weights + bias)
\(\hat{y}_i\) Prediction for example \(i\)
\(y_i\) True target for example \(i\)

Gradient (Level 3 form you will see in notes):

\[\nabla_{\boldsymbol{\theta}} J = \frac{2}{n}\mathbf{X}^\top (\mathbf{X}\boldsymbol{\theta} - \mathbf{y})\]
Symbol Meaning
\(\mathbf{X}^\top\) Transpose of \(\mathbf{X}\) (flips rows/columns)
\((\mathbf{X}\boldsymbol{\theta} - \mathbf{y})\) Vector of residuals (prediction − truth)

Update: \(\boldsymbol{\theta} \leftarrow \boldsymbol{\theta} - \eta \nabla_{\boldsymbol{\theta}} J\).

EXAMPLE (synthetic). True rule \(y = 5x_1 + 2x_2 + 1\) plus small noise. After 80 steps with \(\eta=0.1\), NumPy recovered \(\boldsymbol{\theta} \approx [4.990, 2.001, 0.931]\) — close to \([5, 2, 1]\). Residual MSE ≈ \(0.054\) (noise floor; not zero).

Preview: neural net training#

A neural net stacks many such linear maps with non-linear activations between them. The cost is still a scalar. Backpropagation is the chain rule applied to compute \(\nabla_{\boldsymbol{\theta}} J\) for every parameter. The update rule is still "subtract learning rate times gradient." If you understand this page, the first half of a deep-learning chapter stops looking like magic.

We are not implementing a net here. We are making sure the optimisation story is solid before frameworks hide it.

How It Works (end-to-end story in one place)#

Here is the same pipeline without new notation, so you can rehearse it aloud before an interview.

  1. Represent one example as a feature vector \(\mathbf{x}\). Represent the model as weights \(\mathbf{w}\) and bias \(b\) (or one augmented \(\boldsymbol{\theta}\)).
  2. Predict with a weighted sum: \(\hat{y} = \mathbf{w}\cdot\mathbf{x} + b\). For a batch, use \(\mathbf{X}\mathbf{w}\).
  3. Score the error with a cost such as MSE. The cost is a scalar — one number that summarises how unhappy you are with the current weights.
  4. Differentiate that scalar with respect to every parameter. Bundle those partial derivatives into \(\nabla J\).
  5. Step opposite the gradient: subtract \(\eta\nabla J\). Repeat until the cost flattens or you hit a step budget.
  6. Read the result. If features were meaningful and the model class was rich enough, \(\mathbf{w}\) often has an interpretation (e.g. "each extra study hour adds about 5 points"). If you used a deep net, individual weights are harder to read, but the training loop is still steps 3–5.

That loop is why calculus shows up in ML even when your day job is mostly Python.

Level ladder used in this tutorial#

Level What you saw Example
1 — Intuition Prediction = weighted sum; training = walk downhill on error Exam-score story
2 — Definitions Vector, matrix, norm, derivative, gradient Symbol tables after each equation
3 — Mathematical form MSE gradient \(\frac{2}{n}\mathbf{X}^\top(\mathbf{X}\boldsymbol{\theta}-\mathbf{y})\) Linear regression section
4 — From-scratch code NumPy loops, no scikit-learn maths_for_ml_lab.py Parts 3–5

We deliberately stop before Level-5 "call a library estimator," so the idea is not skipped.

Worked comparison: three learning rates on the noisy regression#

Same Part-5 data, 80 steps, three values of \(\eta\):

\(\eta\) Final MSE (measured) What you should notice
0.001 17.93 Barely moved; weights still near zero
0.1 0.054 Settled near the noise floor; recovered weights ≈ true rule
1.0 24.44 Unstable / poor fit in this short run — not "faster learning"

ASSUMPTION: these numbers are for this synthetic draw (numpy seed 42), not a universal law. The qualitative lesson is stable: tune \(\eta\) by watching the cost curve, not by folklore.

Shape cheat for the regression gradient#

For \(\mathbf{X}\) shaped \((n, d)\), \(\boldsymbol{\theta}\) shaped \((d,)\), \(\mathbf{y}\) shaped \((n,)\):

Expression Shape Role
\(\mathbf{X}\boldsymbol{\theta}\) \((n,)\) Predictions
\(\mathbf{X}\boldsymbol{\theta}-\mathbf{y}\) \((n,)\) Residuals
\(\mathbf{X}^\top(\mathbf{X}\boldsymbol{\theta}-\mathbf{y})\) \((d,)\) Unnormalised gradient direction
\(\frac{2}{n}\mathbf{X}^\top(\cdots)\) \((d,)\) MSE gradient

If a line in your notebook throws matmul shape errors, fix shapes before doubting the formula.

Hands-on lab (45–60 minutes)#

Goal: run and extend maths_for_ml_lab.py so every abstract symbol has a number next to it.

  1. Create the venv (see Setup) and run:
python maths_for_ml_lab.py
  1. Part 1 — confirm prediction 73.0 and batch [73. 83. 60.] by hand.
  2. Part 2 — confirm L2 of \((3,4)\) is \(5\); cosine of parallel vectors is \(1\).
  3. Part 3 — watch \(w\) climb from \(-2\) toward \(3\). Change lr to 1.1 and observe divergence (see Common Mistakes).
  4. Part 4 — note that \(w_2\) reaches \(-1\) faster than \(w_1\) reaches \(2\) (steeper valley).
  5. Part 5 — recover weights near \([5, 2, 1]\). Try lr=0.001 and lr=1.0; record final MSE.

Lab outputs (abridged, real run):

Prediction y_hat = w·x + b = 73.0
Batch predictions X @ w + b = [73. 83. 60.]
||a||_2 = 5.0
cosine_similarity = 1.0000
final 1D: w = 2.981111, J(w) = 0.00035681
final 2D: w = [1.970438, -1.000000], J = 0.00087390
Recovered theta ≈ [4.990, 2.001, 0.931]
View the full lab script (maths_for_ml_lab.py)
#!/usr/bin/env python3
"""
Mathematics for Machine Learning — Hands-on Lab
Companion to: Mathematics for Machine Learning: The Essentials You'll Actually Use
Site: pushpjeet.com/tutorials/mathematics-for-machine-learning/

Run:  python maths_for_ml_lab.py

Needs: Python 3.10+, NumPy. No GPU, no API keys.
Estimated time: 45–60 minutes if you type the exercises yourself.
"""

from __future__ import annotations

import numpy as np

np.set_printoptions(precision=4, suppress=True)


def section(title: str) -> None:
    print("\n" + "=" * 60)
    print(title)
    print("=" * 60)


# ---------------------------------------------------------------------------
# Part 1 — Vectors, matrices, dot product (prediction as weighted sum)
# ---------------------------------------------------------------------------

def part1_vectors_and_dot_product() -> None:
    section("Part 1: Vectors, matrices, and a tiny prediction")

    # EXAMPLE: hours studied (x1) and hours slept (x2) -> predicted exam score
    # Model: y_hat = w1*x1 + w2*x2 + b
    x = np.array([4.0, 7.0])          # features for one student
    w = np.array([8.0, 3.0])          # learned weights
    b = 20.0                          # bias

    y_hat = float(np.dot(w, x) + b)
    print(f"Feature vector x = {x}")
    print(f"Weight vector  w = {w}")
    print(f"Bias           b = {b}")
    print(f"Prediction y_hat = w·x + b = {y_hat}")
    print("(Hand check: 8*4 + 3*7 + 20 = 32 + 21 + 20 = 73)")

    # Batch of 3 students as rows of a matrix X (shape 3 x 2)
    X = np.array([
        [4.0, 7.0],
        [6.0, 5.0],
        [2.0, 8.0],
    ])
    y_batch = X @ w + b              # matrix-vector multiply = weighted sums
    print(f"\nBatch feature matrix X (3 students x 2 features):\n{X}")
    print(f"Batch predictions X @ w + b = {y_batch}")


# ---------------------------------------------------------------------------
# Part 2 — Norms and distance
# ---------------------------------------------------------------------------

def part2_norms_and_distance() -> None:
    section("Part 2: L2 norm and Euclidean distance")

    a = np.array([3.0, 4.0])
    b = np.array([0.0, 0.0])

    l2_a = float(np.linalg.norm(a, ord=2))
    dist = float(np.linalg.norm(a - b, ord=2))
    print(f"a = {a}")
    print(f"||a||_2 = sqrt(3^2 + 4^2) = {l2_a}")
    print(f"distance(a, origin) = {dist}")

    # Cosine similarity (used later in RAG / embeddings)
    u = np.array([1.0, 2.0, 3.0])
    v = np.array([2.0, 4.0, 6.0])  # parallel to u
    cos = float(np.dot(u, v) / (np.linalg.norm(u) * np.linalg.norm(v)))
    print(f"\ncosine_similarity(u, v) for parallel vectors = {cos:.4f} (expect 1.0)")


# ---------------------------------------------------------------------------
# Part 3 — Gradient of a simple quadratic cost (1D)
# ---------------------------------------------------------------------------

def part3_1d_gradient_descent() -> None:
    section("Part 3: Gradient descent on J(w) = (w - 3)^2  [1D]")

    # J(w) = (w - 3)^2
    # dJ/dw = 2(w - 3)
    # Minimum at w* = 3, J* = 0

    def J(w: float) -> float:
        return (w - 3.0) ** 2

    def dJ(w: float) -> float:
        return 2.0 * (w - 3.0)

    w = -2.0
    lr = 0.1
    history = []
    for step in range(25):
        history.append((step, w, J(w), dJ(w)))
        w = w - lr * dJ(w)

    print("step |      w |     J(w) |  dJ/dw")
    print("-----+--------+----------+-------")
    for step, ww, jj, g in history[::5]:  # every 5 steps
        print(f"{step:4d} | {ww:6.3f} | {jj:8.4f} | {g:6.3f}")
    print(f"final: w = {w:.6f}, J(w) = {J(w):.8f}  (target w*=3)")


# ---------------------------------------------------------------------------
# Part 4 — Gradient descent on a 2D quadratic cost
# ---------------------------------------------------------------------------

def part4_2d_gradient_descent() -> dict:
    section("Part 4: Gradient descent on J(w) = (w1-2)^2 + 4*(w2+1)^2  [2D]")

    # Minimum at w* = (2, -1)
    # grad J = [2(w1-2),  8(w2+1)]

    def J(w: np.ndarray) -> float:
        return float((w[0] - 2.0) ** 2 + 4.0 * (w[1] + 1.0) ** 2)

    def grad_J(w: np.ndarray) -> np.ndarray:
        return np.array([2.0 * (w[0] - 2.0), 8.0 * (w[1] + 1.0)])

    w = np.array([0.0, 0.0])
    lr = 0.05
    path = [w.copy()]
    costs = [J(w)]

    for step in range(40):
        g = grad_J(w)
        w = w - lr * g
        path.append(w.copy())
        costs.append(J(w))

    print("step |     w1 |     w2 |    J(w)")
    print("-----+--------+--------+--------")
    for i in [0, 5, 10, 20, 39]:
        ww = path[i]
        print(f"{i:4d} | {ww[0]:6.3f} | {ww[1]:6.3f} | {costs[i]:7.4f}")
    print(f"final: w = [{w[0]:.6f}, {w[1]:.6f}], J = {J(w):.8f}")
    print("target: w* = [2.0, -1.0], J* = 0")

    return {"final_w": w, "final_J": J(w), "path": path, "costs": costs}


# ---------------------------------------------------------------------------
# Part 5 — Tiny linear regression from scratch (2 features + bias)
# ---------------------------------------------------------------------------

def part5_linear_regression_gd() -> None:
    section("Part 5: Linear regression with gradient descent (from scratch)")

    # Synthetic data: y = 5*x1 + 2*x2 + 1 + small noise
    rng = np.random.default_rng(42)
    n = 40
    X_raw = rng.normal(size=(n, 2))
    true_w = np.array([5.0, 2.0])
    true_b = 1.0
    y = X_raw @ true_w + true_b + rng.normal(scale=0.3, size=n)

    # Augment with a column of ones for the bias (design matrix)
    X = np.column_stack([X_raw, np.ones(n)])  # shape (n, 3)
    # Parameters: theta = [w1, w2, b]
    theta = np.zeros(3)
    lr = 0.1
    history = []

    for step in range(80):
        preds = X @ theta
        err = preds - y
        # Mean squared error
        mse = float(np.mean(err ** 2))
        # Gradient of MSE w.r.t. theta: (2/n) X^T (X theta - y)
        grad = (2.0 / n) * (X.T @ err)
        theta = theta - lr * grad
        if step % 20 == 0 or step == 79:
            history.append((step, theta.copy(), mse))

    print("True weights: w=[5, 2], b=1")
    print("step |     w1 |     w2 |      b |     MSE")
    print("-----+--------+--------+--------+--------")
    for step, th, mse in history:
        print(f"{step:4d} | {th[0]:6.3f} | {th[1]:6.3f} | {th[2]:6.3f} | {mse:7.4f}")
    print(f"\nRecovered theta ≈ [{theta[0]:.3f}, {theta[1]:.3f}, {theta[2]:.3f}]")


# ---------------------------------------------------------------------------
# Student exercises (stubs — solutions printed when --solutions is passed)
# ---------------------------------------------------------------------------

def exercise_hints() -> None:
    section("Exercises (do these yourself; answers at end of tutorial)")
    print(
        """
E1. Change the learning rate in Part 3 to 1.1. What happens? Why?
E2. Implement L1 norm of a = [3, -4] by hand and with np.linalg.norm(a, ord=1).
E3. For J(w1,w2) = w1^2 + w2^2, write the gradient by hand, then verify with
    a tiny finite-difference check in NumPy (perturb each coordinate by 1e-5).
E4. In Part 5, try lr=0.001 and lr=1.0. Record final MSE after 80 steps for each.
E5. Explain in one sentence: why do we subtract the gradient (not add it)?
"""
    )


def main() -> None:
    print("Mathematics for Machine Learning — Lab")
    print("Python / NumPy from-scratch examples (no scikit-learn)")
    part1_vectors_and_dot_product()
    part2_norms_and_distance()
    part3_1d_gradient_descent()
    part4_2d_gradient_descent()
    part5_linear_regression_gd()
    exercise_hints()
    print("\nLab finished. Paste key outputs into your notes.")


if __name__ == "__main__":
    main()

Download the lab ZIPCheat sheet PDFCheat sheet (web)

A one-page symbol sheet lives in the companion student cheat sheet (also as a printable PDF).

Common mistakes#

  1. Learning rate too large. On \(J(w)=(w-3)^2\) with \(w_0=-2\) and \(\eta=1.1\), cost grows: \(25 \rightarrow 36 \rightarrow 52 \rightarrow \ldots \rightarrow 321\) by step 7. The step jumps over the minimum and climbs the other side. Fix: reduce \(\eta\) (try \(0.1\), \(0.01\)).

  2. Adding the gradient instead of subtracting. That maximises the cost. Training "works" only if you minimise — always remember the minus sign.

  3. Shape / transpose bugs. X @ w needs matching inner dimensions. Print .shape before debugging maths that is actually fine.

  4. Forgetting the bias. A model forced through the origin cannot fit \(y = 5x + 1\). Augment \(\mathbf{X}\) with a column of ones or keep \(b\) separate and update it with \(\frac{\partial J}{\partial b}\).

  5. Confusing L1 and L2. L2 squares then square-roots; L1 sums absolute values. Regularisation and "distance" mean different geometries.

  6. Treating training loss \(0\) as always possible. With noisy labels (Part 5), MSE bottoms near the noise level (~\(0.05\) here), not zero. Chasing zero then overfits.

  7. Jumping to sklearn.LinearRegression before deriving the gradient once. Libraries are fine in production; they are a poor first teacher. Derive, implement once, then use the library.

Real-world applications#

Area Where this maths appears
Tabular ML Linear / logistic regression, feature scaling (often L2-related), regularisation
Embeddings & RAG Dot products and cosine similarity between vectors — see the RAG tutorial
Neural nets Every dense layer is matrix multiply; training is gradient descent (SGD / Adam variants)
Recommendation User and item vectors; scores as dots
Computer vision / NLP Same linear algebra at huge scale inside convolution and attention (attention scores are scaled dots)

Interview questions#

Q1. What is the difference between a derivative and a gradient?
A: A derivative is the slope of a function of one variable. A gradient is the vector of partial derivatives of a function of several variables. Each component says how the function changes if you nudge that one parameter.

Q2. Why do we subtract the gradient in gradient descent?
A: The gradient points to steepest ascent of the cost. Subtracting it moves parameters toward lower cost. The learning rate scales the step.

Q3. Express a linear prediction with a bias using a dot product.
A: Either \(\hat{y}=\mathbf{w}\cdot\mathbf{x}+b\), or fold \(b\) into an augmented vector: \(\hat{y}=\boldsymbol{\theta}\cdot[x_1,\ldots,x_n,1]\).

Q4. What happens if the learning rate is too large?
A: Updates overshoot the minimum; the loss can oscillate or diverge. (Quote the \(\eta=1.1\) numbers from Common Mistakes if whiteboarding.)

Q5. How is cosine similarity related to the dot product?
A: Cosine is the dot product divided by the product of L2 norms. On unit vectors, cosine equals the dot product — which is why many retrieval stacks normalise embeddings first (RAG tutorial).

MCQs (10)#

1. The expression \(w_1 x_1 + w_2 x_2\) is best described as:

  1. Element-wise product
  2. Dot product
  3. Matrix inverse
  4. Determinant
Show answer

Answer: B

2. In \(\mathbf{w} \leftarrow \mathbf{w} - \eta \nabla J\), \(\eta\) is:

  1. Momentum
  2. Learning rate
  3. Regularisation strength
  4. Batch size
Show answer

Answer: B

3. \(\|\mathbf{a}\|_2\) for \(\mathbf{a}=(3,4)\) equals:

  1. 7
  2. 12
  3. 5
  4. 25
Show answer

Answer: C

4. The gradient \(\nabla J\) points toward:

  1. Steepest descent
  2. Steepest ascent
  3. A random direction
  4. Always the origin
Show answer

Answer: B

5. MSE for regression is:

  1. Mean of absolute residuals
  2. Mean of squared residuals
  3. Max residual
  4. Median residual
Show answer

Answer: B

6. For \(J(w)=(w-3)^2\), the derivative \(dJ/dw\) is:

  1. \(w-3\)
  2. \(2(w-3)\)
  3. \(2w\)
  4. \(3-w\)
Show answer

Answer: B

7. Adding a column of ones to \(\mathbf{X}\) before \(\mathbf{X}\boldsymbol{\theta}\) mainly helps you:

  1. Normalise features
  2. Encode the bias inside \(\boldsymbol{\theta}\)
  3. Compute determinants
  4. Shuffle rows
Show answer

Answer: B

8. Cosine similarity of two parallel non-zero vectors is:

  1. 0
  2. −1
  3. 1
  4. Undefined
Show answer

Answer: C

9. L1 norm of \((3,-4)\) is:

  1. 5
  2. 7
  3. 12
  4. 1
Show answer

Answer: B

10. If training MSE stops near \(0.05\) on noisy data, the most likely explanation is:

  1. Bug in NumPy
  2. Noise floor / irreducible error
  3. Learning rate must be 1.0
  4. Need a deeper net always
Show answer

Answer: B

Practice exercises#

E1. In Part 3 of the lab, set lr = 1.1. Record \(w\) and \(J\) for 8 steps. Explain the growth using the update rule.
(Measured: \(J\) goes \(25 \rightarrow 36 \rightarrow 52 \rightarrow 75 \rightarrow 107 \rightarrow 155 \rightarrow 223 \rightarrow 321\).)

E2. Compute \(\|\mathbf{a}\|_1\) for \(\mathbf{a}=(3,-4)\) by hand and with np.linalg.norm(a, ord=1).
(Both give \(7.0\).)

E3. For \(J(w_1,w_2)=w_1^2+w_2^2\), write \(\nabla J\). At \((1.5,-0.8)\) check with central finite differences (\(\varepsilon=10^{-5}\)).
(Analytic and finite-difference both \([3.0,\,-1.6]\).)

E4. In Part 5, after 80 steps record final MSE for \(\eta\in\{0.001,\,0.1,\,1.0\}\).
(Measured: \(17.93\), \(0.054\), \(24.44\) respectively — too small crawls; too large fails to settle near the optimum in this setup.)

E5. In one sentence: why subtract the gradient?
(Because the gradient points uphill on the cost; we want to go downhill.)

Downloads#

Free companion material for this tutorial. The numbers in all three match the lab run shown on this page (Python 3.13.5, NumPy 2.5.3).

Student cheat sheet (PDF)

Symbols, must-know equations, NumPy patterns and lab checkpoints on one printable sheet.

Download PDF

Hands-on lab (ZIP)

maths_for_ml_lab.py plus a short README: vectors, norms, 1D/2D gradient descent and linear regression from scratch.

Download ZIP

Cheat sheet (web)

The same cheat sheet as a web page, with rendered formulas.

Open cheat sheet

Summary#

  • A linear prediction is a dot product of weights and features, plus bias.
  • Matrices apply that weighted sum to a whole batch at once.
  • L2 norms measure length and distance; cosine is a normalised dot product (RAG retrieval uses this).
  • Derivatives give slope in 1D; gradients package all partial derivatives in higher dimensions.
  • Gradient descent repeatedly walks opposite the gradient, scaled by the learning rate.
  • Linear regression training and (at a high level) neural net training share that optimisation story.
  • Implement the idea once in NumPy. Then frameworks make sense instead of feeling like superstition.

Further learning (primary / well-known sources)#

  • 3Blue1Brown — Essence of linear algebra (visual intuition for vectors, dot products, matrices). Conceptual further learning; watch before heavier texts.
  • Gilbert Strang — Introduction to Linear Algebra (and MIT OCW linear algebra lectures). Classic, clear on the computations ML reuses.
  • Ian Goodfellow, Yoshua Bengio, Aaron Courville — Deep Learning (MIT Press): early chapters on linear algebra, probability, and numerical computation for deep learning.
  • Mathematics for Machine Learning (Deisenroth, Faisal, Ong) — free textbook aimed exactly at this bridge; use as a slower parallel track after this pillar.
  • Pushpjeet — AI & Machine Learning Foundations and RAG Tutorial in Python for where this maths shows up in systems.

FAQ#

Is this enough maths to start deep learning?
Enough to understand loss, gradients, and why training moves weights. You will still need the chain rule / backprop in more detail for deep nets — this page is the on-ramp, not the whole highway.

Do I need to memorise matrix calculus identities?
Memorise the few you use weekly (dot product, MSE gradient for linear regression, gradient descent update). Look up the rest. Understanding beats memorising forty identities.

Why NumPy and not scikit-learn first?
Because LinearRegression().fit hides the gradient. Once you have computed \(\nabla J\) yourself, using scikit-learn is a speed choice, not a conceptual leap.

How does this connect to RAG?
Document and query embeddings are vectors. Ranking by cosine similarity is ranking by a normalised dot product — the same linear algebra as Part 1–2, applied to text. Details: /tutorials/rag-tutorial-python/.

Does this page cover probability, Bayes or softmax?
No. This page stays on the linear algebra and calculus essentials for training. Probability is a separate topic; join the newsletter if you want to hear when new tutorials are published.

Continue

Newsletter

Liked this? Get the next tutorial by email

New tutorials, labs and cheat sheets, plus AWS batch dates and Oracle tips if you want them. No spam, unsubscribe anytime.

I’m interested in

Double opt-in: you’ll get a confirmation email first. Privacy policy