THE GARDEN
Updated 04 Oct 2026
iteratively reduces a differentiable objective by moving opposite its gradient.
θk+1=θk−αk∇J(θk)\theta_{k+1}=\theta_k-\alpha_k\nabla J(\theta_k).
αk\alpha_k is the learning rate or step size.
too large a step can overshoot or diverge; very small steps can converge slowly.
for J(β)=∣Xβ−y∣22/(2n)J(\beta)=|X\beta-y|_2^2/(2n), ∇J(β)=XT(Xβ−y)/n\nabla J(\beta)=X^T(X\beta-y)/n.
Unavailable reference specifies an objective; gradient descent is one way to optimize an objective
the normal equations characterize least-squares stationary solutions without taking gradient-decent steps
◌ Explore connections in Graph view
Paths through the garden