§ 00 · Machine Learning Overview

Formulas learn.
Watch them converge.

Machine learning is not magic — it is calculus finding the path of least error. In this visual guide, gradient descent, neural decision boundaries, cluster groupings, and polynomial regularization emerge live on interactive canvas simulators. Click to add data points, adjust hyperparameters, and watch algorithms update.

class A (green) class B (rose) A live 2-layer Neural Network drawing its decision boundary for a rotating spiral dataset. The background color shows the network's local classification probability.
§ 01 · Linear Regression & Gradient Descent

Tuning sliders of slope and intercept

Linear regression fits a straight line to data by minimizing the **Mean Squared Error (MSE)**. Gradient descent solves this by calculating the local slope of the loss surface and stepping downhill. Set the learning rate α, click to place points, and watch the line seek the middle.

Loss(w,b) = 1/N Σ (w·xᵢ + b − yᵢ)²   —  steps in directions −∂L/∂w and −∂L/∂b
regression line y = wx + b error residuals Click on the canvas above to place new data points. The inset graph plots the Loss Curve over training epochs.
Slope w
–
Intercept b
–
Loss (MSE)
–
Epochs
–

Make learning rate α large (e.g. close to 0.5) and run. The loss may explode as the line overshoots the minimum and bounces out of control. Set α to a small value (e.g. 0.01) and it crawls steadily. This shows why selecting the optimal learning rate is a fundamental tuning step in ML models.

§ 02 · Neural Networks & Classification

Fitted boundaries, bent by activations

A single linear boundary cannot separate complex shapes like concentric rings or XOR quadrants. By feeding linear combinations into nonlinear **activation functions** (like ReLU or Tanh) and stacking layers, a Neural Network bends decision surfaces to wrap around features.

ah = σ(W₁·x + b₁)  →  ŷ = sigmoid(W₂·ah + b₂)
class B (green) class A (rose) Left: Decision boundary heatmap. Right: Live weight connections (teal = positive, rose = negative).
Loss (BCE)
–
Training Accuracy
–%
Epochs
–
Separability
–

Switch the activation function to Linear. No matter how many neurons or epochs you run on the Circle or XOR dataset, the network can only draw a straight line. Without non-linear activations, hidden layers are mathematically useless. Switch to **ReLU** or **Tanh** to see the boundary curve and envelope the clusters.

§ 03 · Unsupervised Clustering

Finding patterns without answers

K-Means is unsupervised: it has no labels or target values. It iteratively moves K centroid stars to minimize the average squared distance to clustered points. Watch the Voronoi cells adjust as centroids take steps toward local density centers.

Inertia (WCSS) = Σ ||xᵢ − μc(i)||²   —  update μₖ to the mean of assigned cluster points
centroid μₖ The shaded partitions display Voronoi boundary cells: points inside each color cell belong to the nearest centroid.
Inertia (WCSS)
–
Current Phase
–
Iterations
–
Status
–

K-Means is sensitive to its **initial centroid placements**. Click "Re-Initialize" multiple times with K = 3 or 4. Sometimes the centroids settle in suboptimal arrangements because the algorithm got stuck in a local minimum. Modern implementations use smart startup heuristics (like K-Means++) to combat this.

§ 04 · Bias-Variance & Regularization

Smoothing the wrinkles of overfitting

Give a model too much capacity (like a high-degree polynomial) and it wiggles through every noisy training point, performing terribly on new test data (high variance). Adding a penalty **λ** to the weight magnitudes (L2 Regularization / Ridge regression) forces the curve to stay smooth.

Cost = MSE + λ·Σ wⱼ²   —  solves via normal equation: w = (XᵀX + λ·I)⁻¹Xᵀy
true sine wave fitted polynomial test points (red dots) Drag any green training point up or down to modify the data. The model computes the optimal polynomial coefficient matrix live.
Train MSE
–
Test MSE
–
Sum of weights (L2)
–
Model State
–

Set polynomial degree to 9 and L2 Penalty to 0. The curve behaves wildly near the edges, hitting training points perfectly but overshooting test markers. Now, drag the L2 slider up to 0.01. The weights shrink dramatically, and the wild curves collapse back into a clean, smooth wave that fits test points. Regularization manages the bias-variance trade-off.

§ 05 · Practice Terminal

Now you compute the learning

Solve these questions using the concepts described above. Hints reveal one step at a time — check your ideas on the interactive simulations!

1 · The learning rate check

A regression model initialized with weight w = 2.0 receives one training point (x = 1.0, y = 3.0) with bias fixed at b = 0. Using Mean Squared Error (MSE) loss and a learning rate α = 0.1, calculate the updated weight w₁ after exactly one gradient descent step.

Hint 1 — Calculate the prediction error: ŷ = w·x = 2(1) = 2. The loss derivative with respect to w for a single sample is: ∂L/∂w = 2 · (ŷ − y) · x.
Hint 2 — Substitute the values: ∂L/∂w = 2 · (2 − 3) · 1 = −2.0. The update rule is w₁ = w − α · (∂L/∂w).
Answer — w₁ = 2.0 − 0.1 · (−2.0) = 2.2. The gradient was negative, pushing the weight upward toward the target of 3.0. Try matching this in § 01 with a single point!

2 · The nonlinear boundary threshold

A single neuron has inputs x = [1, −1], weights w = [1.5, 0.5], and bias b = −0.8. Find the final activation output a if the neuron uses: (a) a ReLU activation, (b) a Sigmoid activation.

Hint 1 — Calculate the pre-activation net input z = w₁·x₁ + w₂·x₂ + b = (1.5)(1) + (0.5)(−1) − 0.8.
Hint 2 — z = 1.5 − 0.5 − 0.8 = 0.2. ReLU output is max(0, z). Sigmoid output is 1 / (1 + e^−z).
Answer — (a) ReLU: a = max(0, 0.2) = 0.2. (b) Sigmoid: a = 1 / (1 + e^−0.2) ≈ 0.550. Since z > 0, both outputs show positive predictions. Verify in § 02 by observing neural node readouts.

3 · L2 weight decay penalty

A model has three weights: w = [2.0, −1.0, 0.5]. The loss without regularization is Loss₀ = 4.2. Calculate the regularized cost if we apply L2 regularization (Ridge) with penalty coefficient λ = 0.5. (Assume the L2 penalty is calculated as λ · Σ wⱼ²).

Hint 1 — Compute the sum of the squared weights: Σ wⱼ² = 2.0² + (−1.0)² + 0.5² = 4.0 + 1.0 + 0.25 = 5.25.
Hint 2 — The regularized cost formula is: Cost = Loss₀ + λ · (Σ wⱼ²). Substitute Loss₀ = 4.2 and λ = 0.5.
Answer — Cost = 4.2 + 0.5 · 5.25 = 4.2 + 2.625 = 6.825. Adding the penalty increased the objective value. Optimization forces weights to shrink to reduce this penalty, smooth the curve, and prevent overfitting. Check it out in § 04.

4 · Diagnosing underfitting vs overfitting

A machine learning engineer trains a classifier. The training dataset loss is 0.02, but the validation loss is 1.45. Is the model underfitting or overfitting, and what is the best immediate solution?

Hint 1 — Compare the training loss and validation loss. A massive gap where training loss is tiny but validation error is massive indicates the model has memorized the noise of the training data.
Hint 2 — This high variance behavior is typical of **overfitting**. To solve it, we need to reduce the model complexity, add regularization, or collect more training data.
Answer — The model is **overfitting** (low bias, high variance). The best solutions are to add an **L2 weight decay penalty (λ)**, reduce complexity (e.g., lower polynomial degree or drop hidden units), or apply dropout. Underfitting would show high loss on both training and validation sets.