🚀NEW LABGetting Started with Claude AgentsStart lab
Safety

Beyond Preference in AI Alignment

First page
Beyond Preference in AI Alignment
Paper summary

challenges the dominant practice of AI alignment known as human preference tuning; explains in what ways human preference tuning fails to capture the thick semantic content of human values; argues that AI alignment needs reframing, instead of aligning on human preferences, AI should align on normative standards appropriate to their social roles.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack