Skip to content

Adam optimizer

The Adam optimizer, short for adaptive moment estimation, is a gradient descent method that gives every parameter its own step size, computed from running averages of recent gradients and their squares.

Adam keeps two exponential moving averages per parameter: a first moment for the mean of past gradients and a second moment for the mean of their squares. Both start at zero, so a bias correction rescales them over the early steps. Each update then divides the corrected first moment by the square root of the corrected second moment, which damps steps along noisy or steep directions.

Below, plain gradient descent and Adam descend a valley that is steep in one parameter and shallow in the other.

Interactive diagram — enable JavaScript to view.

Kingma and Ba introduced Adam in 2014, and it became the standard choice for training deep neural networks because it converges quickly with little tuning. Typical defaults are a learning rate of 0.001 with decay rates of 0.9 and 0.999 for the two moments. The tradeoff is memory, since Adam stores two extra values per weight.

Large language model training runs mostly use AdamW, a variant that applies weight decay directly to the weights instead of routing it through the loss function. That separation generalizes better.

Stochastic Gradient Descent Algorithm With Python and NumPy

Tutorial

Stochastic Gradient Descent Algorithm With Python and NumPy

In this tutorial, you'll learn what the stochastic gradient descent algorithm is, how it works, and how to implement it with Python and NumPy.

advanced algorithms machine-learning numpy

For additional information on related topics, take a look at the following resources:

Have a question about this? Mentor AI can show you examples, compare related terms, and point you to tutorials.


By Martin Breuss • Updated Sept. 16, 2026