Optimizer Laboratory
Race down the loss curves
Every machine learning model is optimized by rolling down a loss surface. This sandbox lets you select a landscape $J(w_1, w_2)$, choose optimizers, set the learning rate, and **click anywhere on the map** to release particles. Watch them race to the minimum.
Loss Surface:
J(w₁, w₂) = 0.5 · (w₁² + 2·w₂²)
SGD
Momentum
RMSprop
Adam
Click anywhere on the contour grid to launch particles for all four optimizers simultaneously. Trace lines show their history.
SGD w
–
Momentum w
–
RMSprop w
–
Adam w
–
Step Count
0
Optimizer Mathematics
SGD Stochastic Gradient Descent |
Updates weights directly in the direction of the negative gradient. Simple, but easily stalls in valleys or oscillates wildly if parameters differ in scale.
wt+1 = wt − α · gt |
Momentum Classical Momentum |
Adds a fraction of the previous step's velocity vector to the update. Simulates a heavy ball rolling with inertia, helping it slide past noisy gradients and escape flat saddle points.
vt = β·vt-1 + α·gt
wt+1 = wt − vt |
RMSprop Root Mean Squared Prop |
Maintains an exponentially decaying average of squared gradients. It divides the learning rate by this running standard deviation, shrinking steps in steep directions and growing them in flat areas.
st = β·st-1 + (1−β)·gt²
wt+1 = wt − [α / √(st + ε)] · gt |
Adam Adaptive Moment Estimation |
Combines the ideas of **Momentum** (first moment) and **RMSprop** (second moment) together, including bias corrections for initialization. The industry standard optimizer for deep neural networks.
mt = β₁·mt-1 + (1−β₁)·gt ; vt = β₂·vt-1 + (1−β₂)·gt²
m̂t = mt / (1−β₁ᵗ) ; v̂t = vt / (1−β₂ᵗ)
wt+1 = wt − [α / √(v̂t + ε)] · m̂t |