University5 min

Gradient Descent

Optimization on loss landscapes

Not Started

Gradient Descent on Loss Landscapes

Click map to set start point

Loss Function

L(x,y) = x² + 2y²

Learning rate α0.050
Momentum β0.000

Status

Position(, )
Loss
‖∇L‖0.0000
Steps-1

x ← x − α∇L(x)

💡 Experiment

Try: 1) Large learning rate → oscillation/divergence. 2) Small lr → slow convergence. 3) Momentum → faster valley traversal. 4) Multi-modal → local minima. Click the map to set a new start.