🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Training · Retrieval

InstructRetro

Free while signed in. Answers cite the passages they came from.

First page
InstructRetro
The curator’s take

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.

Key points
01

48B scale: Continues pretraining a 43B parameter GPT model on 100B additional tokens while retrieving from a 1.2T-token database.

02

Instruction tuning: Further instruction-tunes the retrieval-pretrained model, producing an instruction-following version of Retro.

03

Stronger factuality: Shows reduced hallucination and better factuality on knowledge-intensive tasks compared to Retro-free baselines at comparable scale.

04

Retrieval pretraining validated: Provides evidence that retrieval-during-pretraining can scale to 40B+ parameters and benefit downstream instruction-tuned use cases.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack