Agent-FLAN
Free while signed in. Answers cite the passages they came from.

Agent-FLAN redesigns fine-tuning data so that open models can learn agentic skills without sacrificing general capability, hitting new open-source SoTA for Llama2-7B-based agents.
Format vs reasoning split: The training corpus separates "format following" (producing valid tool/JSON syntax) from "agent reasoning" so each skill can be learned at its own rate and against pre-training distribution.
Negative sampling: Carefully constructed negative examples teach the model when *not* to call a tool, directly targeting hallucinated or spurious actions.
Llama2-7B gains: Agent-FLAN beats prior best open agent fine-tunes by 3.5% averaged across agent evaluation datasets, while preserving general LLM capability.
Scales cleanly: The recipe's improvements hold up as the base model grows, making it a reusable template for future open-source agent fine-tunes.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack