Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas and Cozmin Ududec at the UK AI Security Institute release Transect, an open-source package built on Inspect Scout for analysing very long agent evaluation transcripts in a reproducible way.
Ask this paper
Problem. Long-horizon agentic evaluations produce transcripts of hundreds of pages. A final score hides course corrections and gaming attempts, reviewers cannot read everything, and LLM-assisted analysis gives evaluators many unrecorded analytic choices.
Design. A reusable evaluation-family configuration holds task context and behavioural vocabulary, kept separate from judge models and analysis settings. Reports align events, token use, sub-agent activity and model-generated labels on one turn-based timeline, and every label traces back to its source turns.
Case study. On an AI R&D evaluation of almost 13 million tokens, 89% were spent in delegated sub-agent work and 66.6% of file accesses targeted the manuscript workspace.
Finding. The run concentrated on operational work and manuscript production and showed little evidence of a sustained hypothesis-generation stage.
Reliability checks. Repeated judge rolls give agreement rates per label (81.7% of outputs received unanimous research-activity labels), and the package flags missing or unstable classifications for review.
Abstract
Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.