🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023Issue 182 · Sep 28 – Oct 4, 2026

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,650
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
Paper of the week
AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents
01 · Sep 28 – Oct 4, 2026Code

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

AutoCompact trains a coding agent to decide when to compact its context, what working state to keep, and how to continue afterward, as part of its own policy. A judge reviews the base agent's compaction decisions and replaces flawed ones before they execute, and the corrected trajectories are used for SFT and then for RL that optimizes coding and compaction together on task success. Pass rates rise by 9.2 points on SWE-bench Verified and 5.0 points on SWE-PolyBench Verified, and the gains hold both with a 256K window that never overflows and with a 16K window that falls back to forced compaction.

CodeAgents
This week · 10 papersView the full issue →
Context Language Models

Context Language Models

Agent harnesses usually manage the model's context through fixed rules such as summarization or compaction. Researchers from Meta and collaborators propose Context Language Models (CLMs), which treat the live context as a file the model edits freely, deciding what to keep, rewrite, or remove.

02Agents
Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Salesforce AI Research's Critical-State RL finds the one call in a multi-turn tool-use interaction where training helps, since reward variation that depends on later turns often reflects downstream randomness instead of the current action. It uses nested sampling to separate action-dependent reward variation from continuation noise, then trains only the selected call with contextual-bandit updates. On BFCL v4 missing-function tasks, training the selected turn adds about 14 points, while training the alternative turn leaves accuracy flat or worse.

03Agents
AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

AutoGym: Blueprint-First Generation of Verifiable Agent Gyms

Training agents with RL requires a gym, meaning a task, an executable environment to attempt it in, and a verifier that reliably separates success from failure. These gyms are still built by hand, saturate as models improve, and get exposed to contamination. Researchers from Amazon AGI present AutoGym, which generates complete gyms from a minimal domain seed or from prior model trajectories.

04Agents
SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym: Internalizing Human Skills into LLMs for Real-World Problem Solving

SkillGym turns human-written agent skills into 2,756 verifiable training environments across 12 categories, each with code-based checkers, and collects 8,364 successful trajectories for fine-tuning. Under Claude Code, fine-tuning Qwen3.5-35B-A3B adds 19.10 points on Terminal-Bench 2.1 and 28.13 points on skill-assisted SkillsBench v1.1, where it reaches 51.47%, above the reported scores for Claude Sonnet 4.6 and GPT-5.4 Mini. With no skills loaded, the trained model still beats the base model that has the skills in context.

05Agents
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

Long-term memory and self-improvement are usually built as separate systems around the agent loop. Researchers from MIT CSAIL built JAZ to test how far a minimal harness, little more than the agent loop itself, can go on the tasks those systems target.

06Agents
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

On long-horizon tasks, decisions such as which hypothesis to test or which implementation to build on determine how the whole run turns out. Researchers from Microsoft and collaborators call the ability to make these decisions well an agent's taste, and build Taste-Bench to measure it.

07Agents
Coding Agents are Strong Prompt Optimizers

Coding Agents are Strong Prompt Optimizers

Search-based prompt optimizers such as GEPA propose edits, run fresh rollouts, score them, and keep only the edits that improve a validation metric. Researchers from Microsoft show that this loop may be unnecessary when you already have a corpus of agent trajectories.

08Agents
Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Multi-agent harnesses usually depend on a central orchestrator that assigns tasks and coordinates workers, and that orchestrator limits how many agents the system can use. Microsoft Research introduces Agensh, a self-organized multi-agent harness with no central orchestrator, and scales it to 1,024 coding agents.

09Agents
Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Many agentic memory systems use an autoregressive LLM to decide how memories are organized, retrieved, and used, which puts expensive generation on the critical path of every memory operation. Jev-Mem borrows its design from System-One/System-Two cognition and hands those decisions to a lightweight controller.

10Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026