🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation · Safety

Evaluation Feature Steering in LLMs

Paper preview
Evaluation Feature Steering in LLMs
Paper summary

evaluates featuring steering in LLMs using an experiment that artificially dials up and down various features to analyze changes in model outputs; it focused on 29 features related to social biases and study if feature steering can help mitigate social biases; among its findings, it reports that feature steering sometimes leads to off-target effects and that a neutrality feature can help decreases social biases in 9 social dimensions without negatively affecting text quality.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack