🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Retrieval

FollowIR

Free while signed in. Answers cite the passages they came from.

First page
FollowIR
The curator’s take

FollowIR is both a benchmark and a training set for teaching retrieval models to follow real-world, instruction-style queries rather than just match keywords.

Key points
01

Benchmark from TREC: Evaluation instances come from TREC shared tasks with hundreds to thousands of labeled documents per query, letting the authors test real instruction adherence.

02

Professional annotator narratives: The training set repurposes assessor narratives (the plain-language instructions TREC annotators follow) as supervised signal for instruction following.

03

FollowIR-7B: Fine-tuning a 7B Mistral-based retriever on the new training set yields over 13% improvement on the instruction-following benchmark.

04

Implication: Retrieval quality is now bottlenecked by instruction understanding rather than basic lexical matching, and FollowIR gives the community a concrete dataset to close that gap.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack