🚀NEW LABGetting Started with Claude AgentsStart lab
← All papersIssue 181 of 182

The week of Sep 21 – Sep 27, 2026

10 papers, hand-picked and summarised.

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

In agent workloads, a short tool call can return a long search result or execution trace that has to be prefilled before decoding resumes, and the context keeps growing across turns. Xiaomi's MiMo team built HySparse2, the attention architecture behind the upcoming MiMo-V3, to lower prefill cost and KV-cache size while improving long-context retrieval.

01Efficiency
Harness-Zero: Harness Distillation via Agent-as-Harness

Harness-Zero: Harness Distillation via Agent-as-Harness

A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

02Agents
XYEval: Agents say yes to bad advice

XYEval: Agents say yes to bad advice

Users often suggest a fix that sounds right and is wrong, and Google DeepMind's XYEval measures how often agents go along with it by adding one confident, misleading hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE, and MCP-Atlas while keeping the correct solution unchanged. Scores fall by up to 46.7% relative across Gemini, Claude Opus 4.8, and GPT 5.5, and agents often disagree with the hint in their reasoning and then follow it without telling the user. A system prompt warning about the XY problem helps on single-turn tasks but leaves large drops on multi-turn ones such as tau2-bench and SWE-bench Verified.

03Agents
JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

Running a frontier LLM as the judge on every eval gets expensive at scale. This paper tests a cheaper setup, where a decision-only judge handles most calls and only the uncertain ones go to a frontier model.

04Evaluation
Self-Organizing Agent Teams Learn to Reason Together

Self-Organizing Agent Teams Learn to Reason Together

Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

05Agents
WFM: Wiki Foundation Model for Complex Agentic Reasoning

WFM: Wiki Foundation Model for Complex Agentic Reasoning

More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

06Agents
EvoOntology: A Self-Evolving Ontology Layer for Data Agents

EvoOntology: A Self-Evolving Ontology Layer for Data Agents

EvoOntology replaces the hand-written semantic layer that data agents usually get in their prompt with an ontology they query at runtime, built by a dedicated builder agent and served over MCP with schema, content, and tool layers. The ontology evolves through small typed edits, and each edit is kept only if a paired evaluation on the same backbone shows it helps. On DDR-Bench, accuracy rises 17.8 points on average across backbones, and on BIRD, execution accuracy rises 7.4 points, with tool-layer edits accounting for 57% of the gain from evolution.

07Agents
Self Improvement via Fast Tree-search

Self Improvement via Fast Tree-search

Coding agents that rewrite their own implementation can improve on benchmarks, but prior methods such as the Darwin Gödel Machine (DGM) are expensive to run. Researchers from MIT and Sakana AI trace most of that cost to one step and make it cheaper.

08Evaluation
GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

LLM plans for long-horizon robot tasks often break embodiment constraints, fail to recover from mistakes, or lose track of objects they cannot see. GAVEL adds an explicit graph world model around the LLM and more than doubles the success rate of a small model without changing its weights.

09Agents
ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

Google Cloud AI Research built ScientistTwo, a multi-agent framework that takes a problem from a human expert and runs the full discovery cycle without further intervention, from establishing baselines and screening ideas on a data subset to running its own ablations and revising the idea from them. Manuscript drafting includes a simulated peer-review and rebuttal engine. Benchmarked on problems from papers accepted at ICLR, ICML, and NeurIPS, its solutions outperform the human state-of-the-art models, and its papers score higher average ratings than the human-authored ones under automated AI reviewers.

10Agents
Every Monday
Get next week’s papers.
Subscribe on Substack