RAG vs. Finetuning

Microsoft researchers systematically compare RAG and fine-tuning (and their combination) on LLMs like Llama 2 and GPT-4 using an agricultural domain dataset.
Ask this paper
Domain-specific testbed: Agricultural Q&A is chosen precisely because it highlights gaps in LLMs' parametric knowledge about specialized, region-specific domains.
Fine-tuning helps: Fine-tuning alone lifts accuracy by over 6 percentage points versus the base model - non-trivial but not a full solution.
RAG helps more, and stacks: RAG adds another ~5 percentage points on top of fine-tuning, and the gains are cumulative, suggesting the two techniques target different failure modes.
Practitioner playbook: Argues that real-world domain LLM deployments should view RAG and fine-tuning as complementary, not alternatives - a guideline that has since hardened into common practice.