🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Sep 27, 2026
Agents

Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior

First page
Who Is Behind the Harness? Fingerprinting LLMs through Agentic Behavior
The curator’s take

Chuyi Wang, Yong Cui and colleagues at Tsinghua University present LIDAR, a black-box method that identifies which LLM runs behind a coding-agent harness from its actions rather than its text.

Ask this paper

Key points
01

Probe pairs. Three coding probe pairs expose post-edit verification, recovery from transient failures, and how the agent resolves conflicts between specification and tests.

02

Features. Trajectories are represented with instance-level and distribution-level features and compared with clean references by a lightweight probabilistic identifier; no weights or logits are needed.

03

Results. Across 36 models from seven families, LIDAR reaches 95.13% and 88.36% Top-1 accuracy (MRR 0.9736 and 0.9336) under two harnesses and beats four fingerprinting and API-auditing baselines.

04

Why it matters. Swapping the model behind an agent changes security-relevant behavior, and this gives operators a way to check what they are running.

Abstract

LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer. We present LIDAR (LLM Identification from Decisions and Actions at Runtime), an active black-box fingerprinting method for coding-agent execution. Three coding probe pairs expose post-edit verification, transient-failure recovery, and specification-test conflict resolution under controlled changes. LIDAR represents the resulting trajectories with complementary instance-level and distribution-level features and compares them with clean references using a lightweight probabilistic identifier. It requires no access to model weights, logits, or provider internals. Across 36 models from seven families and two agent harnesses, LIDAR achieves high Top-1 accuracy and MRR and outperforms four existing fingerprinting and API-auditing baselines. Ablations confirm that the two feature levels, all probe pairs, and their controlled variants contribute. These results show that agent execution behavior provides model-identity evidence beyond final outputs.

Every Monday
Get next week’s papers.
Subscribe on Substack