🚀NEW LABGetting Started with Claude AgentsStart lab
Training · Retrieval

InstructRetro

First page
InstructRetro
Paper summary

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

Ask this paper

Key points
01

48B scale: Continues pretraining a 43B parameter GPT model on 100B additional tokens while retrieving from a 1.2T-token database.

02

Instruction tuning: Further instruction-tunes the retrieval-pretrained model, producing an instruction-following version of Retro.

03

Stronger factuality: Shows reduced hallucination and better factuality on knowledge-intensive tasks compared to Retro-free baselines at comparable scale.

04

Retrieval pretraining validated: Provides evidence that retrieval-during-pretraining can scale to 40B+ parameters and benefit downstream instruction-tuned use cases.

Every Monday
Get next week’s papers.
Subscribe on Substack