Skip to content

Latest commit

 

History

History
105 lines (105 loc) · 3.23 KB

2022-06-28-borgeaud22a.md

File metadata and controls

105 lines (105 loc) · 3.23 KB
title booktitle abstract layout series publisher issn id month tex_title firstpage lastpage page order cycles bibtex_author author date address container-title volume genre issued pdf extras
Improving Language Models by Retrieving from Trillions of Tokens
Proceedings of the 39th International Conference on Machine Learning
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
inproceedings
Proceedings of Machine Learning Research
PMLR
2640-3498
borgeaud22a
0
Improving Language Models by Retrieving from Trillions of Tokens
2206
2240
2206-2240
2206
false
Borgeaud, Sebastian and Mensch, Arthur and Hoffmann, Jordan and Cai, Trevor and Rutherford, Eliza and Millican, Katie and Van Den Driessche, George Bm and Lespiau, Jean-Baptiste and Damoc, Bogdan and Clark, Aidan and De Las Casas, Diego and Guy, Aurelia and Menick, Jacob and Ring, Roman and Hennigan, Tom and Huang, Saffron and Maggiore, Loren and Jones, Chris and Cassirer, Albin and Brock, Andy and Paganini, Michela and Irving, Geoffrey and Vinyals, Oriol and Osindero, Simon and Simonyan, Karen and Rae, Jack and Elsen, Erich and Sifre, Laurent
given family
Sebastian
Borgeaud
given family
Arthur
Mensch
given family
Jordan
Hoffmann
given family
Trevor
Cai
given family
Eliza
Rutherford
given family
Katie
Millican
given family
George Bm
Van Den Driessche
given family
Jean-Baptiste
Lespiau
given family
Bogdan
Damoc
given family
Aidan
Clark
given family
Diego
De Las Casas
given family
Aurelia
Guy
given family
Jacob
Menick
given family
Roman
Ring
given family
Tom
Hennigan
given family
Saffron
Huang
given family
Loren
Maggiore
given family
Chris
Jones
given family
Albin
Cassirer
given family
Andy
Brock
given family
Michela
Paganini
given family
Geoffrey
Irving
given family
Oriol
Vinyals
given family
Simon
Osindero
given family
Karen
Simonyan
given family
Jack
Rae
given family
Erich
Elsen
given family
Laurent
Sifre
2022-06-28
Proceedings of the 39th International Conference on Machine Learning
162
inproceedings
date-parts
2022
6
28