🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Thinking LLMs

First page
Thinking LLMs
Paper summary

proposes a training method to equip LLMs with thinking abilities for general instruction-following without human-annotated data; uses an iterative search and optimization procedure to explore thought generation which enables the model to learn without direct supervision; thought candidates for each user instruction are scored with a judge model; only responses are evaluated by the Judge which determines the best and worst ones; then the corresponding full outputs are used as chosen and rejected pairs for DPO (referred to as Thought Preference Optimization in this paper). reports superior performance on AlpacaEval and Arena-Hard.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack