Optimizer Laboratory

Race down the loss curves

Every machine learning model is optimized by rolling down a loss surface. This sandbox lets you select a landscape $J(w_1, w_2)$, choose optimizers, set the learning rate, and **click anywhere on the map** to release particles. Watch them race to the minimum.

Loss Surface:
J(w₁, w₂) = 0.5 · (w₁² + 2·w₂²)
SGD Momentum RMSprop Adam Click anywhere on the contour grid to launch particles for all four optimizers simultaneously. Trace lines show their history.
SGD w
–
Momentum w
–
RMSprop w
–
Adam w
–
Step Count
0

Optimizer Mathematics

SGD
Stochastic Gradient Descent
Updates weights directly in the direction of the negative gradient. Simple, but easily stalls in valleys or oscillates wildly if parameters differ in scale.
wt+1 = wt − α · gt
Momentum
Classical Momentum
Adds a fraction of the previous step's velocity vector to the update. Simulates a heavy ball rolling with inertia, helping it slide past noisy gradients and escape flat saddle points.
vt = β·vt-1 + α·gt
wt+1 = wt − vt
RMSprop
Root Mean Squared Prop
Maintains an exponentially decaying average of squared gradients. It divides the learning rate by this running standard deviation, shrinking steps in steep directions and growing them in flat areas.
st = β·st-1 + (1−β)·gt²
wt+1 = wt − [α / √(st + ε)] · gt
Adam
Adaptive Moment Estimation
Combines the ideas of **Momentum** (first moment) and **RMSprop** (second moment) together, including bias corrections for initialization. The industry standard optimizer for deep neural networks.
mt = β₁·mt-1 + (1−β₁)·gt  ;  vt = β₂·vt-1 + (1−β₂)·gt²
m̂t = mt / (1−β₁ᵗ)  ;  v̂t = vt / (1−β₂ᵗ)
wt+1 = wt − [α / √(v̂t + ε)] · m̂t