🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 19, 2026
Evaluation

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes

First page
SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes
The curator’s take

Mengxiao Wang and Nitesh Saxena present FARSIGHT, a scheme-level evaluation of financial LLM trading agents on robustness under market turbulence and security against three attack classes, and apply it to 15 academic schemes.

Ask this paper

Key points
01

80 percent fail a core robustness metric. Under flash-crash-like scenarios, which most of the surveyed schemes never test.

02

100 percent show security vulnerabilities. Across attacks on information sources, attacks on agents, and agent-as-attacker behaviour.

03

The two failures are linked. A small misjudgement can cascade into a market-wide crash by itself, and an adversary can trigger the same collapse deliberately at low cost, so robustness and security cannot be evaluated separately here.

04

The scope is a systematization. 118 references and 15 schemes, which makes it a reference point for what the published trading-agent literature does and does not check.

Abstract

Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.

Every Monday
Get next week’s papers.
Subscribe on Substack