Et Tu, Brute? Economic Misalignment in Personal AI Agents

Aman Priyanshu and Supriti Vijay (Foundation AI, Cisco) with Brian Jabarian and Niloofar Mireshghallah (Carnegie Mellon) show that personal AI agents given a user's inbox or profile recommend more expensive options to users they infer are wealthy, even when the request is identical.
Ask this paper
Scale. 325K experiments across 13 models from four families and three domains (flights, health insurance, graduate programs), each with a controlled catalog of 200 items. Eight of the 13 models show the effect.
Size of the gap. Wealth-conditioned differences reach $198 per flight and $284 per month for insurance with Claude Opus 4.8, and almost $3,900 per year for graduate programs. The shift is asymmetric: wealthy users are moved up $85 on flights while low-income users move down $51.
Overrides the user. When told to find the cheapest flight, Gemini 2.5 Flash still recommends options $208 more expensive for wealthy personas.
Masking fails. Blocking financial attributes largely removes the gap, but blocking other attributes leaves it or increases it (GPT-5.5 insurance gap up 40% to $151 when employment is hidden). Wealth inferred from unrelated emails preserves much of the gap.
Capability does not help. Larger models are no better, and Claude Opus 4.8 shows the largest effect. The authors name the pattern adversarial delegation.
Abstract
Personal AI agents make recommendations and take actions on people's behalf in high-stakes economic contexts, e.g., buying a flight, choosing health insurance, or selecting a graduate program. The agent is given access to the user's personal context, e.g., their email inbox and a structured profile of personal attributes, with the intention of making an optimal, personalized decision for the user. We show that by simply providing this personal context, the agent steers recommendations based on inferred wealth, without being explicitly instructed to do so. In a suite of 325K experiments on 13 agents across three types of economic decisions (flights, health insurance, and graduate programs), we find that 8 models systematically choose more expensive options for wealthier users when requests are identical. This steering continues even when it directly goes against the user's stated objective: when explicitly instructed to find the cheapest option, some agents still act on the wealth profile they have inferred. It also occurs when wealth is inferred from ambient data, such as emails unrelated to the task. And it persists under privacy controls that block specific attributes: blocking financial attributes largely removes the disparity, but blocking other attributes leaves it unchanged and can increase it by up to 40% for insurance, as agents rely on the remaining signals to infer wealth. Larger and more capable models are no better; Claude Opus 4.8 shows the largest effect. We term this misalignment "adversarial delegation", in which the very conditions that make a personal AI agent useful - access to personal information - enable it to act against the user's interests.