Formulas learn.
Watch them converge.
Machine learning is not magic — it is calculus finding the path of least error. In this visual guide, gradient descent, neural decision boundaries, cluster groupings, and polynomial regularization emerge live on interactive canvas simulators. Click to add data points, adjust hyperparameters, and watch algorithms update.
Tuning sliders of slope and intercept
Linear regression fits a straight line to data by minimizing the **Mean Squared Error (MSE)**. Gradient descent solves this by calculating the local slope of the loss surface and stepping downhill. Set the learning rate α, click to place points, and watch the line seek the middle.
Make learning rate α large (e.g. close to 0.5) and run. The loss may explode as the line overshoots the minimum and bounces out of control. Set α to a small value (e.g. 0.01) and it crawls steadily. This shows why selecting the optimal learning rate is a fundamental tuning step in ML models.
Fitted boundaries, bent by activations
A single linear boundary cannot separate complex shapes like concentric rings or XOR quadrants. By feeding linear combinations into nonlinear **activation functions** (like ReLU or Tanh) and stacking layers, a Neural Network bends decision surfaces to wrap around features.
Switch the activation function to Linear. No matter how many neurons or epochs you run on the Circle or XOR dataset, the network can only draw a straight line. Without non-linear activations, hidden layers are mathematically useless. Switch to **ReLU** or **Tanh** to see the boundary curve and envelope the clusters.
Finding patterns without answers
K-Means is unsupervised: it has no labels or target values. It iteratively moves K centroid stars to minimize the average squared distance to clustered points. Watch the Voronoi cells adjust as centroids take steps toward local density centers.
K-Means is sensitive to its **initial centroid placements**. Click "Re-Initialize" multiple times with K = 3 or 4. Sometimes the centroids settle in suboptimal arrangements because the algorithm got stuck in a local minimum. Modern implementations use smart startup heuristics (like K-Means++) to combat this.
Smoothing the wrinkles of overfitting
Give a model too much capacity (like a high-degree polynomial) and it wiggles through every noisy training point, performing terribly on new test data (high variance). Adding a penalty **λ** to the weight magnitudes (L2 Regularization / Ridge regression) forces the curve to stay smooth.
Set polynomial degree to 9 and L2 Penalty to 0. The curve behaves wildly near the edges, hitting training points perfectly but overshooting test markers. Now, drag the L2 slider up to 0.01. The weights shrink dramatically, and the wild curves collapse back into a clean, smooth wave that fits test points. Regularization manages the bias-variance trade-off.
Now you compute the learning
Solve these questions using the concepts described above. Hints reveal one step at a time — check your ideas on the interactive simulations!
1 · The learning rate check
A regression model initialized with weight w = 2.0 receives one training point (x = 1.0, y = 3.0) with bias fixed at b = 0. Using Mean Squared Error (MSE) loss and a learning rate α = 0.1, calculate the updated weight w₁ after exactly one gradient descent step.
ŷ = w·x = 2(1) = 2. The loss derivative with respect to w for a single sample is: ∂L/∂w = 2 · (ŷ − y) · x.∂L/∂w = 2 · (2 − 3) · 1 = −2.0. The update rule is w₁ = w − α · (∂L/∂w).w₁ = 2.0 − 0.1 · (−2.0) = 2.2. The gradient was negative, pushing the weight upward toward the target of 3.0. Try matching this in § 01 with a single point!2 · The nonlinear boundary threshold
A single neuron has inputs x = [1, −1], weights w = [1.5, 0.5], and bias b = −0.8. Find the final activation output a if the neuron uses: (a) a ReLU activation, (b) a Sigmoid activation.
z = w₁·x₁ + w₂·x₂ + b = (1.5)(1) + (0.5)(−1) − 0.8.z = 1.5 − 0.5 − 0.8 = 0.2. ReLU output is max(0, z). Sigmoid output is 1 / (1 + e^−z).a = max(0, 0.2) = 0.2. (b) Sigmoid: a = 1 / (1 + e^−0.2) ≈ 0.550. Since z > 0, both outputs show positive predictions. Verify in § 02 by observing neural node readouts.3 · L2 weight decay penalty
A model has three weights: w = [2.0, −1.0, 0.5]. The loss without regularization is Loss₀ = 4.2. Calculate the regularized cost if we apply L2 regularization (Ridge) with penalty coefficient λ = 0.5. (Assume the L2 penalty is calculated as λ · Σ wⱼ²).
Σ wⱼ² = 2.0² + (−1.0)² + 0.5² = 4.0 + 1.0 + 0.25 = 5.25.Cost = Loss₀ + λ · (Σ wⱼ²). Substitute Loss₀ = 4.2 and λ = 0.5.Cost = 4.2 + 0.5 · 5.25 = 4.2 + 2.625 = 6.825. Adding the penalty increased the objective value. Optimization forces weights to shrink to reduce this penalty, smooth the curve, and prevent overfitting. Check it out in § 04.4 · Diagnosing underfitting vs overfitting
A machine learning engineer trains a classifier. The training dataset loss is 0.02, but the validation loss is 1.45. Is the model underfitting or overfitting, and what is the best immediate solution?