🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning · Data

SaySelf

First page
SaySelf
Paper summary

a training framework to teach LLMs to express more accurate fine-grained confidence estimates and self-reflective rationales; it performs supervised finetuning on a dataset that contains summaries of the difference between multiple reasoning chains; reinforcement learning is then applied to calibrate confidence estimates, encouraging the LLM to produce accurate, high-confidence predictions and penalize overconfidence in erroneous outputs.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack