🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning

rStar

First page
rStar
Paper summary

introduces self-play mutual reasoning to improve the reasoning capabilities of small language models without fine-tuning or superior models; MCTS is augmented with human-like reasoning actions, obtained from SLMs, to build richer reasoning trajectories; a separate SLM provides unsupervised feedback on the trajectories and the target SLM selects the final reasoning trajectory as the answer; rStar boosts GSM8K accuracy from 12.51% to 63.91% for LLaMA2-7B and consistently improves the accuracy of other SLMs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack