Paper
Scaling Laws for Neural Language Models (2020): loss falls as a power law in parameters, data and compute
Kaplan, McCandlish et al. (OpenAI) measure how language-model loss scales smoothly with model size, dataset size and compute across seven orders of magnitude, and derive the compute-optimal allocation. The paper turned 'bigger is better' into a budget: it justified GPT-3 and the capex race that followed. 30 pages.