title

booktitle

abstract

layout

series

publisher

issn

id

month

tex_title

firstpage

lastpage

page

order

cycles

bibtex_author

author

date

address

container-title

volume

genre

issued

pdf

extras

Improving Language Models by Retrieving from Trillions of Tokens

Proceedings of the 39th International Conference on Machine Learning

We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a 2 trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25{\texttimes} fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.

inproceedings

Proceedings of Machine Learning Research

PMLR

2640-3498

borgeaud22a

0

Improving Language Models by Retrieving from Trillions of Tokens

2206

2240

2206-2240

2206

false

Borgeaud, Sebastian and Mensch, Arthur and Hoffmann, Jordan and Cai, Trevor and Rutherford, Eliza and Millican, Katie and Van Den Driessche, George Bm and Lespiau, Jean-Baptiste and Damoc, Bogdan and Clark, Aidan and De Las Casas, Diego and Guy, Aurelia and Menick, Jacob and Ring, Roman and Hennigan, Tom and Huang, Saffron and Maggiore, Loren and Jones, Chris and Cassirer, Albin and Brock, Andy and Paganini, Michela and Irving, Geoffrey and Vinyals, Oriol and Osindero, Simon and Simonyan, Karen and Rae, Jack and Elsen, Erich and Sifre, Laurent

given	family
Sebastian	Borgeaud

given	family
Arthur	Mensch

given	family
Jordan	Hoffmann

given	family
Trevor	Cai

given	family
Eliza	Rutherford

given	family
Katie	Millican

given	family
George Bm	Van Den Driessche

given	family
Jean-Baptiste	Lespiau

given	family
Bogdan	Damoc

given	family
Aidan	Clark

given	family
Diego	De Las Casas

given	family
Aurelia	Guy

given	family
Jacob	Menick

given	family
Roman	Ring

given	family
Tom	Hennigan

given	family
Saffron	Huang

given	family
Loren	Maggiore

given	family
Chris	Jones

given	family
Albin	Cassirer

given	family
Andy	Brock

given	family
Michela	Paganini

given	family
Geoffrey	Irving

given	family
Oriol	Vinyals

given	family
Simon	Osindero

given	family
Karen	Simonyan

given	family
Jack	Rae

given	family
Erich	Elsen

given	family
Laurent	Sifre

2022-06-28

Proceedings of the 39th International Conference on Machine Learning

162

inproceedings

date-parts

2022

6

28

https://proceedings.mlr.press/v162/borgeaud22a/borgeaud22a.pdf

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

2022-06-28-borgeaud22a.md

2022-06-28-borgeaud22a.md

Files

2022-06-28-borgeaud22a.md

Latest commit

History

2022-06-28-borgeaud22a.md

File metadata and controls