AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

Context Language Models
Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove.

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms
Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use
Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing
In agent workloads, a short tool call can return a long search result or execution trace that has to be prefilled before decoding resumes, and the context keeps growing across turns. Xiaomi's MiMo team built HySparse2, the attention architecture behind the upcoming MiMo-V3, to lower prefill cost and KV-cache size while improving long-context retrieval.

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving
SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it.

Coding Agents are Strong Prompt Optimizers
Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.

Agensh: Scaling Organizational Intelligence to 1,024 Agents
Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.

XYEval: Agents say yes to bad advice
Users often suggest a fix that sounds right and is wrong, and Google DeepMind's XYEval measures how often agents go along with it by adding one confident, misleading hint to tasks from tau2-bench, SWE-bench, Terminal-Bench, HLE, and MCP-Atlas while keeping the correct solution unchanged. Scores fall by up to 46.7% relative across Gemini, Claude Opus 4.8, and GPT 5.5, and agents often disagree with the hint in their reasoning and then follow it without telling the user. A system prompt warning about the XY problem helps on single-turn tasks but leaves large drops on multi-turn ones such as tau2-bench and SWE-bench Verified.

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents
Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

Harness-Zero: Harness Distillation via Agent-as-Harness
A specialized harness can raise an agent's performance a lot, but the best harness differs across domains, instances, and models. Harness-Zero, from Google and colleagues, uses the specialized harness only during training and moves the behavior it induces into the model weights.

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
Running a frontier LLM as the judge on every eval gets expensive at scale. This paper tests a cheaper setup, where a decision-only judge handles most calls and only the uncertain ones go to a frontier model.

Self-Organizing Agent Teams Learn to Reason Together
Multi-agent systems usually fix roles and protocols in advance. Researchers from Stanford and Together AI let a fixed team of models learn how to organize its own collaboration from past exchanges.

WFM: Wiki Foundation Model for Complex Agentic Reasoning
More agents now store long-term memory as an LLM Wiki, a folder of markdown pages linked to each other. Each page holds dense text and the links hold structure, and WFM is a Wiki Foundation Model trained to use both when retrieving.

EvoOntology: A Self-Evolving Ontology Layer for Data Agents
EvoOntology replaces the hand-written semantic layer that data agents usually get in their prompt with an ontology they query at runtime, built by a dedicated builder agent and served over MCP with schema, content, and tool layers. The ontology evolves through small typed edits, and each edit is kept only if a paired evaluation on the same backbone shows it helps. On DDR-Bench, accuracy rises 17.8 points on average across backbones, and on BIRD, execution accuracy rises 7.4 points, with tool-layer edits accounting for 57% of the gain from evolution.

Self Improvement via Fast Tree-search
Coding agents that rewrite their own implementation can improve on benchmarks, but prior methods such as the Darwin Gödel Machine (DGM) are expensive to run. Researchers from MIT and Sakana AI trace most of that cost to one step and make it cheaper.

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning
LLM plans for long-horizon robot tasks often break embodiment constraints, fail to recover from mistakes, or lose track of objects they cannot see. GAVEL adds an explicit graph world model around the LLM and more than doubles the success rate of a small model without changing its weights.

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
Agent harnesses are tuned by hand, one mechanism at a time, against whatever environment the team happens to have. NVIDIA moves that tuning into an automated research loop and keeps only the mechanisms that survive selection across many environments.

ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
Google Cloud AI Research built ScientistTwo, a multi-agent framework that takes a problem from a human expert and runs the full discovery cycle without further intervention, from establishing baselines and screening ideas on a data subset to running its own ablations and revising the idea from them. Manuscript drafting includes a simulated peer-review and rebuttal engine. Benchmarked on problems from papers accepted at ICLR, ICML, and NeurIPS, its solutions outperform the human state-of-the-art models, and its papers score higher average ratings than the human-authored ones under automated AI reviewers.

Skill-based Agentic Evaluation for Real-time Data Science Tasks
Storing a fixed reference answer for every eval case breaks when the underlying data changes daily, so Adobe researchers write each reference answer as a Python function that runs against the live system at evaluation time. An LLM judge then splits the agent's response and the computed answer into atomic facts and scores precision and recall regardless of output format, raising agreement with expert labels from an MCC of 0.331 to 0.427 while cutting token cost per case by 16%. A judge given no ground truth scored an MCC of -0.379, which is worse than chance.

Verifiable Social Reasoning for LLM Assistants
People ask assistants for social advice constantly, and the assistant only hears the user's version of events, which makes it hard to check whether it read the situation correctly. Google Research builds that ground truth by simulation, with a target agent holding a hidden motive while a user agent relays events to the assistant, which then has to infer the motive. Across 24k human annotations validating the simulations and 12 LLMs tested, biased framing from the user shifted the assistant's answer, and longer conversations with room for clarifying questions did not reliably help.

Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems
NVIDIA compared eight strategies for choosing which models go into a multi-agent system, based on size, accuracy, answer diversity, and error diversity, across routing, majority vote, and LLM-as-judge setups on hard science benchmarks. Larger pools of different open models raised the theoretical best-case accuracy while achieved accuracy often fell below the single best model in the pool, and using several copies of one model worked better. Majority vote over the best single model raised HLE accuracy from 29.4% to 32.2%, so measure what another model adds before putting it in the router.