🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Evaluation · Data

LongWriter

First page
LongWriter
Paper summary

proposes AgentWrite to enable off-the-shelf LLMs to generate coherent outputs beyond 20K words; AgentWrite breaks down the long generation task into subtasks and in a divide-and-conquer approach generates; the agent breaks the task into multiple writing subtasks and concatenates the outputs to get a final output (i.e., plan + write); the approach is then used to build SFT datasets that are used to tune LLMs to generate coherent longer outputs automatically; a 9B parameter model, further improved through DPO, achieves state-of-the-art performance on their benchmark, and surpasses proprietary models.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack