Paper
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (2018)
Devlin, Chang, Lee and Toutanova (Google AI Language) pre-train a bidirectional Transformer with masked-language modeling and next-sentence prediction, then fine-tune it to top eleven NLP benchmarks at once. BERT made pre-train-then-fine-tune the standard recipe and powered Google Search. 16 pages.
Companies in this report