InstructRetro
Free while signed in. Answers cite the passages they came from.

NVIDIA introduces Retro 48B, the largest LLM pretrained with retrieval at the time.
48B scale: Continues pretraining a 43B parameter GPT model on 100B additional tokens while retrieving from a 1.2T-token database.
Instruction tuning: Further instruction-tunes the retrieval-pretrained model, producing an instruction-following version of Retro.
Stronger factuality: Shows reduced hallucination and better factuality on knowledge-intensive tasks compared to Retro-free baselines at comparable scale.
Retrieval pretraining validated: Provides evidence that retrieval-during-pretraining can scale to 40B+ parameters and benefit downstream instruction-tuned use cases.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack