🚀NEW LABGetting Started with Claude AgentsStart lab
Retrieval · Training

RAG vs. Finetuning

First page
RAG vs. Finetuning
Paper summary

Microsoft researchers systematically compare RAG and fine-tuning (and their combination) on LLMs like Llama 2 and GPT-4 using an agricultural domain dataset.

Ask this paper

Key points
01

Domain-specific testbed: Agricultural Q&A is chosen precisely because it highlights gaps in LLMs' parametric knowledge about specialized, region-specific domains.

02

Fine-tuning helps: Fine-tuning alone lifts accuracy by over 6 percentage points versus the base model - non-trivial but not a full solution.

03

RAG helps more, and stacks: RAG adds another ~5 percentage points on top of fine-tuning, and the gains are cumulative, suggesting the two techniques target different failure modes.

04

Practitioner playbook: Argues that real-world domain LLM deployments should view RAG and fine-tuning as complementary, not alternatives - a guideline that has since hardened into common practice.

Every Monday
Get next week’s papers.
Subscribe on Substack