Paper
Attention Is All You Need (2017): the Transformer paper that started the modern AI era
Vaswani et al. (Google Brain/Research) replace recurrence and convolution with self-attention: the Transformer trains in parallel, scales with compute, and reached state-of-the-art translation quality with a fraction of the training cost. Every GPT, Claude, Gemini and Llama descends from this architecture. 15 pages.