Paper
Language Models are Few-Shot Learners (GPT-3, 2020): 175B parameters and the discovery of in-context learning
Brown et al. (OpenAI) scale an autoregressive Transformer to 175 billion parameters and show it performs tasks from a prompt and a few examples with no gradient updates. GPT-3's few-shot ability is the capability that made prompting a product and set the stage for ChatGPT. 75 pages.