InstructRetro
First page

Paper summary
NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.
Ask this paper
01
48B scale: Continues pretraining a 43B parameter GPT model on 100B additional tokens while retrieving from a 1.2T-token database.
02
Instruction tuning: Further instruction-tunes the retrieval-pretrained model, producing an instruction-following version of Retro.
03
Stronger factuality: Shows reduced hallucination and better factuality on knowledge-intensive tasks compared to Retro-free baselines at comparable scale.
04
Retrieval pretraining validated: Provides evidence that retrieval-during-pretraining can scale to 40B+ parameters and benefit downstream instruction-tuned use cases.