Survey on Instruction Tuning for LLMs
First page

Paper summary
A comprehensive survey of instruction tuning covering methodology, dataset construction, and applications.
Ask this paper
01
Systematic literature review: Provides a structured taxonomy of instruction-tuning research across datasets, training recipes, and evaluation approaches.
02
Dataset construction: Reviews how instruction datasets are assembled - from human-written prompts to model-generated self-instruct and hybrid pipelines.
03
Training methodologies: Catalogs SFT, multitask learning, RLHF, and their variants, with a focus on how each technique interacts with instruction-tuning data.
04
Open problems: Highlights issues including instruction-data quality, data scaling, multilingual instruction tuning, and evaluation of instruction-following reliability.