🚀NEW LABGetting Started with Claude AgentsStart lab
Reasoning · Training · Evaluation

RATIONALYST

First page
RATIONALYST
Paper summary

a model for process-supervision of reasoning that enables generalization across diverse reasoning tasks; this process is achieved with pre-training on a collection of 79k rationales from the Pile and a combination of reasoning datasets with minimal human intervention; fine-tuned from LLaMa-3-8B, the proposed model improves the accuracy of reasoning by an average of 3.9% on 7 reasoning benchmarks.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack