Gradient Descent Optimizer

Iteratively minimize f(x) by following the negative gradient

Parameters

ⓘ
ⓘ
ⓘ
ⓘ
Show Trail

Controls

xⓘ

Calculated Values

Final x:
0.00;0.00;
f(x) at minimum:
0.00;0.00;
Steps taken:
40.00;40.00;
Gradient |f'|:
0.00;0.00;

Examples

Parabola x² from x₀ = 2

Minimum at x = 0 with η = 0.1.

  • x: 0.000.00
  • f: 0.000.00

Double well x⁴ − 3x²

Two minima; depends on initial guess.

  • x: 0.000.00
  • f: 0.000.00

Visualization

Gradient Descent in Physics & ML

Gradient descent (steepest descent) is a first-order iterative optimization algorithm for finding a local minimum of a differentiable function. The negative gradient −∇f points in the direction of steepest decrease.

At each step: x_{n+1} = x_n − η·f′(x_n), where η is the learning rate. The animation traces the path from x₀ to the minimum, showing each intermediate evaluation.

In physics, gradient descent appears in energy minimization (relaxation methods), variational calculations, and fitting models to data by minimizing χ² or loss functions.

The learning rate η controls convergence: too small → slow; too large → oscillation or divergence. Adaptive methods like Adam adjust η per step in machine learning.

Key Concepts

  • Negative gradient direction = steepest descent
  • Learning rate η sets step size
  • Only finds local minima (depends on x₀)
  • Converges slowly near flat regions
  • Can oscillate if η is too large

Real-World Applications

  • Energy minimization in molecular dynamics
  • Neural network training (backpropagation)
  • Variational quantum Monte Carlo
  • Least-squares fitting of physics models

Explore Further

More computational physics tools

  • 1D Heat Equation

    Finite-difference FTCS solution to the diffusion equation with animated temperature profiles.

  • 1D Wave Equation

    Leapfrog finite-difference solution to the wave equation with animated wave propagation.

  • Numerical Integration

    Trapezoidal and Simpson rules to approximate definite integrals with error vs exact solutions.

  • ODE Solver

    Euler and Runge-Kutta 4 methods for first-order ODEs with comparison to analytic solutions.

  • Monte Carlo Intro

    Estimate π and integrals by random sampling — introduction to stochastic computational physics.

  • Newton-Raphson

    Solve nonlinear equations f(x) = 0 with tangent-line iterations — fast when the guess is good.

Physics Equations

Update rule:
xn+1=xn−η⋅f′(xn)x_{n+1} = x_n - \eta \cdot f'(x_n)
Converged when:
∣f′(xn)∣<ε|f'(x_n)| < \varepsilon

Step-by-Step Solution

See how the main results are calculated.

1

Step 1: Objective function

Minimise f(x) = x² starting from x₀ = 2.

Calculation:

f(x0)=4.0000,f′(x0)=4.0000f(x_0) = 4.0000,\quad f'(x_0) = 4.0000
2

Step 2: Update rule

Learning rate η = 0.1.

Equation:

xn+1=xn−η f′(xn)x_{n+1} = x_n - \eta\, f'(x_n)

Explanation:

Each step moves opposite to the gradient, i.e. downhill toward a minimum.

3

Step 3: First iteration

Calculation:

x1=2−0.1×4.0000=1.6000x_1 = 2 - 0.1\times 4.0000 = 1.6000

Result:

f(x1)=2.5600f(x_1) = 2.5600
4

Step 4: After 40 steps

Calculation:

x≈0.000266,f(x)≈0.000000x \approx 0.000266,\quad f(x) \approx 0.000000

Result:

∣f′(x)∣=5.32e−4|f'(x)| = 5.32e-4

Explanation:

Gradient is near zero — converged to a minimum.

Frequently Asked Questions (FAQ)

What happens with large learning rate?

The optimizer may overshoot and oscillate around the minimum or diverge entirely.

Can it find global minima?

Not guaranteed. Try multiple starting points or methods like simulated annealing.

Practice MCQs

  1. Gradient descent moves in the direction of:
  2. Too large η causes:
  3. For f(x) = x², the minimum is at:
  4. Gradient descent is a ___ order method.
  5. Local minima problem means:
  6. Momentum in gradient descent helps:
  7. For f(x) = x⁴ − 3x², how many local minima exist?
  8. A convex function guarantees:
  9. If gradient descent oscillates, the likely fix is:
  10. Stochastic gradient descent (SGD) differs from full GD by: