AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence
Ante Kapetanovic and colleagues run 192,000 evaluations to show that putting a prior score in a judge's context metadata drags its rating toward that number, breaking the independence assumption every refinement pipeline relies on.

Where vs What: Decomposing Structural and Content Failures in LLM-Generated Structured Outputs
Yiwei Zhang, Chengke Wu, Li Wang and Jianqiang Li split structured-output failures into placement errors and value errors and find that structure breaks down well before content does.

Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
Lezhi Yu and colleagues (Zhejiang University) name a failure mode in LLM research agents that execution-based benchmarks cannot see: methodological hallucination, where the code runs and the conclusion is still fabricated.

Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay
Prateek Chhikara introduces matched trajectory replay, a protocol that holds answer states, evidence, budgets, and action costs fixed so confidence-to-action mappings in retrieval agents can be compared on their trajectory-level consequences rather than in isolation.

Agent Seer: Synthesizing Scenarios from Specification Understanding
Harish Karumuri, Mahesh Vemula, and David Lopes Pegna show that an MCP specification alone carries enough semantic information to synthesize realistic multi-turn agent evaluation scenarios, with no examples, no live tool access, and no domain tuning.

Evaluating Language Models in Realistic Conversational Contexts
Ilija Subasic, Andrew Rabinovich, and Zhao Chen (Upwork) release UPHELD, a reference-full benchmark of professionally scripted human-to-human dialogues with 36,000 per-turn human annotations, then show standard automatic metrics and LLM judges correlate poorly with expert judgment.

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Kuan-Hao Tseng and colleagues (University of Sydney) build FaulT-Bench, 200 network troubleshooting scenarios across eight topologies that include false fault reports and wrong root-cause claims, then show SADE, ReAct, and Claude Code all collapse when the network is actually healthy.

Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
Leonardo Liparulo and Francesco Pierri (Politecnico di Milano) build an MCP server that mirrors a proprietary hardware design tool and benchmark seven locally deployed open-source models on dependency-ordered engineering workflows, isolating which harness choices actually move reliability.

Same Model, Different Harness: Different Coding-Agent Results
Sydney Lewis holds the model and task fixed and varies only the harness, showing that a coding agent's benchmark number is a property of the model-plus-harness pair rather than of the weights.

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Recuris splits agent memory in two, with a Working Memory tracking task progress and an Experiential Memory holding skills, so skill selection is grounded in the current task state rather than the full growing history. Because skill use is anchored to an explicit state, a failed run points at a specific memory component, and a fixed Meta-Agent turns that evidence into validation-gated updates to Skill Memory. It improves task success in 35 of 37 completed model-benchmark pairs, adding 17.8 points to GPT-5.6 Sol on tau-bench and taking Claude Opus 5 to 87.9%.

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
FM-Bench turns football club management into a 20-year test of sustained agent decision-making. Fifteen frontier models operate through 26 tools and hundreds of consequential decisions in a deterministic environment with no LLM judge. The results show that model scale, price, vendor, and token spend do not predict performance; managerial behavior and memory discipline do. Every model also fails to learn hidden market prices from repeated feedback, exposing a concrete limit in long-horizon adaptation.

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workflows, and recurring failure modes. However, existing memory benchmarks for agents mostly focus on user histories, short traces, or downstream task success, leaving open how to directly evaluate whether memory systems effectively internalize environment-specific experience. To address this gap, we introduce LongMemEval-V2 (LME-V2), a benchmark for evaluating whether memory systems can help agents acquire the experience needed to become knowledgeable colleagues in customized environments. LME-V2 contains 451 manually curated questions covering five core memory abilities for web agents: static state recall, dynamic state tracking, workflow knowledge, environment gotchas, and premise awareness. Questions are paired with history trajectories containing up to 500 trajectories and 115M tokens. We use a context gathering formulation: memory systems consume history trajectories and return compact evidence for downstream question answering. We propose a suite of two memory methods: AgentRunbook-R, an efficient RAG-based memory with knowledge pools for raw state observations, events, and strategy notes, and AgentRunbook-C, which stores trajectories as files and invokes a coding agent to gather evidence in an augmented sandbox. Experiments show that AgentRunbook-C achieves the best performance with 72.5% average accuracy, outperforming the strongest RAG baseline (48.5%) and the off-the-shelf coding agent baseline (69.3%). Despite the strong performance gains, coding agent based methods have high latency costs. While AgentRunbook-C advances the accuracy-latency Pareto frontier, substantial room for improvement remains. Together, these results establish LME-V2 as a challenging testbed for developing long-term memory systems for environment experience.

Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
Shows that an agent's benchmark score predicts its descendants' improvement poorly, names that the Metaproductivity-Performance Mismatch, and offers a measurable substitute aggregated over a clade. It answers which variant a self-modification search should expand next.

Measuring AI Ability to Complete Long Software Tasks
The trend line the talk opens on. Measuring capability as the length of task a system completes, rather than a single-turn score, is what makes harness progress visible at all: the static-harness era and the self-improving era are two slopes on this chart.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Open-ended machine learning research environments scored for both AI agents and human experts under matched time budgets. It is the measurement behind the question the labs are actually asking, which is whether AI can do AI research.

DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
The first entry where the harness stops being hand-written. You cannot backpropagate through a prompt, so DSPy searches over prompts against a small train set instead, and the system prompt becomes an optimised artifact rather than an author's guess.

Judges as a Lifecycle
Most teams validate an LLM judge once, ship it, and never look at it again. Netflix runs judges over hundreds of thousands of show-level recommendation explanations per week, served to millions of members on mobile, and this writeup describes what it takes to keep one honest at that volume.

Skill Lift
Enterprise teams reviewing shared skill libraries almost always gate on a scanner that checks structure, style, and security. NVIDIA measured whether that gate predicts anything about how a skill actually performs, and the answer is close to no.

Meta^n
Systems that edit themselves have to leave part of their own editing machinery untouched to stay stable, which caps realized meta-depth at roughly two. Meta^n keeps the meta-operation fixed and recurses on its input instead, applying one operator repeatedly to its own products and letting convergence set the depth rather than fixing it in advance. Across two backbones it outperforms prior self-improving agents on all eight benchmark families, and on ARC-AGI-2 it is the only method scoring above zero.

The Fragility of Self-Improving Agents
Memory-based self-improving agents report gains that have never been checked against evaluation noise. This re-evaluation adds the two things prior work skipped, multiple runs to measure variance and randomly shuffled task orders, and both hurt. Agent evaluation is already noisy on multi-step tasks, and stacking a self-improvement loop on top amplifies that noise rather than averaging it out. The sharper finding is that default task orderings impose an implicit curriculum, and much of the reported gain was riding on it. Adding detailed rubrics and environment feedback to memory construction recovers part of the drop, and a significant gap remains. If you are measuring your own memory loop, shuffle the task order first.

The Bitter Lesson of Tool Calling
Tool calling is a design choice and the default choice is JSON. For code-capable models, exposing tools as code instead lets calls chain and parallelize naturally, but nobody had run the comparison on an established benchmark across model generations under realistic conditions.

Harness-IF
When a coding agent obeys your rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell compliance from coincidence because they concentrate rules in the user turn, while coding-agent benchmarks only score final task success.

Lost in Compaction
Context compaction is now standard in long-running agent systems, and it silently drops the instructions users most expect to persist. This work names that class, Session Constraints, instructions like "do not delete any emails until I confirm" meant to bind behavior for the rest of a session, and introduces COMPINT to evaluate compactors across multi-turn chat, agentic trajectory, and long-horizon research. Current compactors retain only 17% of injected constraints on average, and most leave the task worse off than running it without compaction at all. Retention swings with the compactor, the prompt, the context length, the phrasing, and where the constraint was injected, which is what makes the loss structural rather than a quirk of one setup. The fix is small and does not touch the compactor or the model: an SC-aware extractor running alongside as a plug-and-play module recovers over 90% retention in all three scenarios.

Model or Harness
Agent evaluations mostly report system-level outcomes, so a failed run leaves the repair unassigned. The same visible failure might call for model post-training, harness engineering, environment redesign, or benchmark repair, and outcome labels cannot separate those cases.