Optimization for Data Science · Interactive demo

Gradient descent and conditioning

The same starting point, the same algorithm, and the same step size. Change the curvature to see why the condition number controls convergence.

Fixed smoothness L = 1Step size η = 1/L = 1Minimizer x⋆ = (0, 0)
κ =
of 456

Changing κ restarts both runs. Click either contour plot to move the shared starting point.

Perfectly conditioned

κ = 1

μ = 1  ·  L = 1  ·  circular level sets

Gradient descent on the perfectly conditioned quadraticCircular contours with the shared initial point and the gradient descent path. With step size one, this quadratic reaches the origin in one step.
Relative objective 1.000
Current point (3.000, −2.000)

Ill-conditioned

κ = 50

μ = 0.0200  ·  L = 1  ·  elliptical level sets

Gradient descent on the adjustable quadraticElliptical contours with the same starting point and step size. Progress slows along the direction with smaller curvature.
Relative objective 1.000
Current point (3.000, −2.000)

Open circle: start  ·  filled dot: current iterate  ·  cross: minimizer. In slice view, a dashed line shows the selected slice. Both plots use the same spatial scale; contour levels are chosen separately for visibility.

Objective convergence

κ = 1: actual = boundRight: actualRight: theorem bound
Normalized objective and the strongly convex gradient descent boundThe logarithmic plot compares f(x_t) divided by f(x_0) with the upper bound (1 minus 1 over kappa) to the power t. Exact zeros use a separate row.

Vertical scale: f(xt) / f(x0). Values below 10⁻⁸ sit at the axis floor; exact zeros use a separate row. The dashed curve shows the full theoretical bound.

Quadratic upper and lower bounds

Three surfaces over (y₁, y₂), tangent at the current iterate. Drag to rotate; use the checkboxes to inspect each surface.

Three-dimensional quadratic surfaces. The function lies between the lower and upper models shown in the formulas below.

Horizontal axes: y₁ and y₂. Vertical axis: objective / model value. All surfaces share the same axes, held fixed while GD runs. The marked point is (xt, f(xt)).

Lower: qμ(y) = f(xt) + ∇f(xt)ᵀ(y − xt) + ½μ‖y − xt‖²

Upper: qL(y) = f(xt) + ∇f(xt)ᵀ(y − xt) + ½L‖y − xt‖²

qμ(y) ≤ f(y) ≤ qL(y) for every y. The bounds follow the selected run’s current iterate; a lower bound can be negative even though f ≥ 0.

The quadratic, the guarantee, and the starting point

In coordinates rotated by 30°, the right-hand objective is

fκ(x) = ½(u²/κ + v²),   (u, v) = Rᵀx.

Its Hessian has eigenvalues μ = 1/κ and L = 1. The left-hand objective is ½‖x‖². Increasing κ lowers the curvature in one direction while preserving the same smoothness constant and step size.

For smooth, strongly convex functions, gradient descent with η = 1/L guarantees

f(xt) / f(x0) ≤ (1 − 1/κ)t.

Here f⋆ = 0. These particular quadratics can converge faster than the bound: one step eliminates the L-eigendirection, and the remaining objective contracts by (1 − 1/κ)² each step. At κ = 1, the entire error disappears in one step.

Shared start:

Coordinates are limited to [−4, 4]. At the minimizer, both runs stay at zero and the relative-objective ratio is undefined. The timeline ends when the general bound reaches 10⁻⁴; playback speed changes only the animation.