Mathematics for Machine Learning — Lab
Python / NumPy from-scratch examples (no scikit-learn)

============================================================
Part 1: Vectors, matrices, and a tiny prediction
============================================================
Feature vector x = [4. 7.]
Weight vector  w = [8. 3.]
Bias           b = 20.0
Prediction y_hat = w·x + b = 73.0
(Hand check: 8*4 + 3*7 + 20 = 32 + 21 + 20 = 73)

Batch feature matrix X (3 students x 2 features):
[[4. 7.]
 [6. 5.]
 [2. 8.]]
Batch predictions X @ w + b = [73. 83. 60.]

============================================================
Part 2: L2 norm and Euclidean distance
============================================================
a = [3. 4.]
||a||_2 = sqrt(3^2 + 4^2) = 5.0
distance(a, origin) = 5.0

cosine_similarity(u, v) for parallel vectors = 1.0000 (expect 1.0)

============================================================
Part 3: Gradient descent on J(w) = (w - 3)^2  [1D]
============================================================
step |      w |     J(w) |  dJ/dw
-----+--------+----------+-------
   0 | -2.000 |  25.0000 | -10.000
   5 |  1.362 |   2.6844 | -3.277
  10 |  2.463 |   0.2882 | -1.074
  15 |  2.824 |   0.0309 | -0.352
  20 |  2.942 |   0.0033 | -0.115
final: w = 2.981111, J(w) = 0.00035681  (target w*=3)

============================================================
Part 4: Gradient descent on J(w) = (w1-2)^2 + 4*(w2+1)^2  [2D]
============================================================
step |     w1 |     w2 |    J(w)
-----+--------+--------+--------
   0 |  0.000 |  0.000 |  8.0000
   5 |  0.819 | -0.922 |  1.4189
  10 |  1.303 | -0.994 |  0.4865
  20 |  1.757 | -1.000 |  0.0591
  39 |  1.967 | -1.000 |  0.0011
final: w = [1.970438, -1.000000], J = 0.00087390
target: w* = [2.0, -1.0], J* = 0

============================================================
Part 5: Linear regression with gradient descent (from scratch)
============================================================
True weights: w=[5, 2], b=1
step |     w1 |     w2 |      b |     MSE
-----+--------+--------+--------+--------
   0 |  0.741 |  0.339 |  0.230 | 22.9987
  20 |  4.797 |  1.999 |  0.943 |  0.0892
  40 |  4.978 |  2.011 |  0.932 |  0.0545
  60 |  4.989 |  2.003 |  0.931 |  0.0543
  79 |  4.990 |  2.001 |  0.931 |  0.0543

Recovered theta ≈ [4.990, 2.001, 0.931]

============================================================
Exercises (do these yourself; answers at end of tutorial)
============================================================

E1. Change the learning rate in Part 3 to 1.1. What happens? Why?
E2. Implement L1 norm of a = [3, -4] by hand and with np.linalg.norm(a, ord=1).
E3. For J(w1,w2) = w1^2 + w2^2, write the gradient by hand, then verify with
    a tiny finite-difference check in NumPy (perturb each coordinate by 1e-5).
E4. In Part 5, try lr=0.001 and lr=1.0. Record final MSE after 80 steps for each.
E5. Explain in one sentence: why do we subtract the gradient (not add it)?


Lab finished. Paste key outputs into your notes.
