🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 8, 2026
Agents · Evaluation

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

First page
Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations
The curator’s take

Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.

Ask this paper

Key points
01

Problem. Long-horizon agentic evaluations produce transcripts of hundreds of pages. A final score hides course corrections and gaming attempts, reviewers cannot read everything, and LLM-assisted analysis gives evaluators many unrecorded analytic choices.

02

Design. A reusable evaluation-family configuration holds task context and behavioural vocabulary, kept separate from judge models and analysis settings. Reports align events, token use, sub-agent activity and model-generated labels on one turn-based timeline, and every label traces back to its source turns.

03

Case study. On an AI R&D evaluation of almost 13 million tokens, 89% were spent in delegated sub-agent work and 66.6% of file accesses targeted the manuscript workspace.

04

Finding. The run concentrated on operational work and manuscript production and showed little evidence of a sustained hypothesis-generation stage.

05

Reliability checks. Repeated judge rolls give agreement rates per label (81.7% of outputs received unanimous research-activity labels), and the package flags missing or unstable classifications for review.

Abstract

Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.

Every Monday
Get next week’s papers.
Subscribe on Substack