🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 3, 2026
Safety

Representational alignment yields generalizable safety in language models

First page
Representational alignment yields generalizable safety in language models
The curator’s take

Lingyu Li, Yan Teng, Yingchun Wang and Xia Hu show that behavioral alignment learns the right answers while leaving the underlying moral category structure untouched, and that fixing the representation instead buys adversarial robustness.

Ask this paper

Key points
01

Prototype theory as the diagnostic: human moral concepts organize around central cases with graded typicality. Across 23 LLMs, models often fail to separate opposed moral categories or preserve typicality within one.

02

The deficit is not a scale or stage artifact: it persists across parameter sizes and across alignment stages, which rules out the usual it-will-fix-itself response.

03

Representational similarity optimization: aligns latent representations with categorization from human moral judgements directly, without supervising any generated response.

04

A clean matched comparison: on the same 251,334 moral annotations, behavioral alignment learned the intended judgements at the response level, left categorization structure largely unchanged, and increased adversarial vulnerability.

05

The trade is explicit: reorganizing categorization gives more modest gains on explicit judgements but consistently improves adversarial robustness across model scales, benchmarks and attack strategies.

Abstract

Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.

Every Monday
Get next week’s papers.
Subscribe on Substack