🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Agents · Training

Agent-FLAN

Free while signed in. Answers cite the passages they came from.

First page
Agent-FLAN
The curator’s take

Agent-FLAN redesigns fine-tuning data so that open models can learn agentic skills without sacrificing general capability, hitting new open-source SoTA for Llama2-7B-based agents.

Key points
01

Format vs reasoning split: The training corpus separates "format following" (producing valid tool/JSON syntax) from "agent reasoning" so each skill can be learned at its own rate and against pre-training distribution.

02

Negative sampling: Carefully constructed negative examples teach the model when *not* to call a tool, directly targeting hallucinated or spurious actions.

03

Llama2-7B gains: Agent-FLAN beats prior best open agent fine-tunes by 3.5% averaged across agent evaluation datasets, while preserving general LLM capability.

04

Scales cleanly: The recipe's improvements hold up as the base model grows, making it a reusable template for future open-source agent fine-tunes.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack