🚀NEW LABGetting Started with Claude AgentsStart lab
Reinforcement Learning

Aligning LLMs with Demonstrated Feedback

First page
Aligning LLMs with Demonstrated Feedback
Paper summary

proposes a method to align LLMs to a specific setting via a very small number of demonstrations as feedback; it aligns LLM outputs to a user’s demonstrated behaviors and can learn fine-grained style and task alignment across domains; outperforms few-shot prompting, SFT, and self-play methods on the tested benchmarks.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack