Paper
Adam: A Method for Stochastic Optimization (2014): the default optimizer for nearly every neural network since
Kingma and Ba combine momentum and per-parameter adaptive learning rates with bias correction into an optimizer that is cheap, robust to hyperparameters and works across vision, language and RL. Adam (and AdamW) remains the default for training LLMs a decade on. 15 pages.