🚀NEW LABGetting Started with Claude AgentsStart lab
DAIR.AI · Curated weekly since April 2023

AI Papers of the Week

Every paper worth reading in AI, hand-picked one week at a time.

2,668
Papers
182
Weekly issues
2023
Since

Discover and explore top AI papers with Claude Code or Codex

npx @dair-ai/mcp setup
1127 papers · AgentsClear filters →
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

Jiayu Shi and Luzhuo Chen release Paritok-4B, a 4B LoRA compressor that shrinks coding-agent context to a quarter of its size by extracting spans rather than paraphrasing them, with weights and data open under Apache 2.0.

601Efficiency
EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Shared skill libraries are usually treated as a safe way for coding agents to reuse each other's work, and EvoMal shows they propagate malware. A planted malicious skill is never invoked, but the agent retrieves it as an authoring template, writes a new skill that preserves the payload, and each authored copy re-enters the library to be imitated again. Across six models the self-poisoning rate runs 20.3% to 41.8%, deleting every planted skill does not clean it up, and a counter-prompt discouraging banner-style copying drops it to 6.7%.

602Code
Tunable Tool-Call Rates in LLM Agents via Representation Steering

Tunable Tool-Call Rates in LLM Agents via Representation Steering

Yuqi Chen, Vincent Siu, Dawn Song and Chenguang Wang show that whether an instruction-tuned model calls a tool is controlled by a single linear direction in its residual stream, extractable without training and steerable at inference.

603Agents
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recuris splits agent memory in two, with a Working Memory tracking task progress and an Experiential Memory holding skills, so skill selection is grounded in the current task state rather than the full growing history. Because skill use is anchored to an explicit state, a failed run points at a specific memory component, and a fixed Meta-Agent turns that evidence into validation-gated updates to Skill Memory. It improves task success in 35 of 37 completed model-benchmark pairs, adding 17.8 points to GPT-5.6 Sol on tau-bench and taking Claude Opus 5 to 87.9%.

604Agents
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

FM-Bench turns football club management into a 20-year test of sustained agent decision-making. Fifteen frontier models operate through 26 tools and hundreds of consequential decisions in a deterministic environment with no LLM judge. The results show that model scale, price, vendor, and token spend do not predict performance; managerial behavior and memory discipline do. Every model also fails to learn hidden market prices from repeated feedback, exposing a concrete limit in long-horizon adaptation.

605Agents
OpenJarvis: Personal AI, On Personal Devices

OpenJarvis: Personal AI, On Personal Devices

The same argument pointed at your laptop. Decompose the personal AI stack into five primitives, then let a frontier cloud model search over that spec while everything runs locally at inference. Roughly 800x lower marginal cost, with the harness rather than the model closing the gap.

606Agents
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues

Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workflows, and recurring failure modes. However, existing memory benchmarks for agents mostly focus on user histories, short traces, or downstream task success, leaving open how to directly evaluate whether memory systems effectively internalize environment-specific experience. To address this gap, we introduce LongMemEval-V2 (LME-V2), a benchmark for evaluating whether memory systems can help agents acquire the experience needed to become knowledgeable colleagues in customized environments. LME-V2 contains 451 manually curated questions covering five core memory abilities for web agents: static state recall, dynamic state tracking, workflow knowledge, environment gotchas, and premise awareness. Questions are paired with history trajectories containing up to 500 trajectories and 115M tokens. We use a context gathering formulation: memory systems consume history trajectories and return compact evidence for downstream question answering. We propose a suite of two memory methods: AgentRunbook-R, an efficient RAG-based memory with knowledge pools for raw state observations, events, and strategy notes, and AgentRunbook-C, which stores trajectories as files and invokes a coding agent to gather evidence in an augmented sandbox. Experiments show that AgentRunbook-C achieves the best performance with 72.5% average accuracy, outperforming the strongest RAG baseline (48.5%) and the off-the-shelf coding agent baseline (69.3%). Despite the strong performance gains, coding agent based methods have high latency costs. While AgentRunbook-C advances the accuracy-latency Pareto frontier, substantial room for improvement remains. Together, these results establish LME-V2 as a challenging testbed for developing long-term memory systems for environment experience.

607Evaluation
Continual Harness: Online Adaptation for Self-Improving Foundation Agents

Continual Harness: Online Adaptation for Self-Improving Foundation Agents

The last step before online learning. Continual Harness keeps history, memory, skills, prompts, and sub-agent specs across trajectories and mutates them while the agent runs, then goes further and updates the weights DAgger-style from what just happened. The presenter calls test-time training the direction that matters most.

608Agents
PostTrainBench: Can LLM Agents Automate LLM Post-Training?

PostTrainBench: Can LLM Agents Automate LLM Post-Training?

Hands an agent one base model, one GPU and ten hours, then asks it to post-train the model on its own. It measures AI automating AI training more directly than any other benchmark here, and audits the reward hacking that follows.

609Training
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing

Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing

Evolves a group of agents that share experience across branches instead of a tree of isolated variants, and reports the strongest coding results in this lineage. The most convincing successor to the Darwin Gödel Machine so far.

610Agents
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

Shows that an agent's benchmark score predicts its descendants' improvement poorly, names that the Metaproductivity-Performance Mismatch, and offers a measurable substitute aggregated over a clade. It answers which variant a self-modification search should expand next.

611Agents
Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents

Self-evolution drifts into unsafe behavior along four paths, which are the model, its memory, its tools and its workflow. The first broad safety study of self-evolving agents, and the counterweight to everything above it.

612Agents
AlphaEvolve: A coding agent for scientific and algorithmic discovery

AlphaEvolve: A coding agent for scientific and algorithmic discovery

An evolutionary coding agent that discovered new algorithms and optimized Google's own infrastructure. It accelerated training of the model that powers it, which makes it the clearest published case of a deployed system improving its own production pipeline.

613Code
A Self-Improving Coding Agent

A Self-Improving Coding Agent

A coding agent with ordinary tools edits its own codebase and improves from 17% to 53% on a random subset of SWE-bench Verified, with no gradient updates. The simplest demonstration that one agent can be both the system under repair and the engineer doing it.

614Code
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Open-ended machine learning research environments scored for both AI agents and human experts under matched time budgets. It is the measurement behind the question the labs are actually asking, which is whether AI can do AI research.

615Agents
Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement

Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement

An agent that reads and rewrites its own logic at runtime, including the logic it uses to rewrite itself, guided only by a high-level objective. It is the first language-model framework to implement the Gödel machine structure directly, with evaluation in place of proof.

616Agents
Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

Spawning another agent becomes just another tool call. The framing here, agents with roles that persist and can be addressed, is what makes sub-agents an addressable resource rather than a one-shot fan-out.

617Agents
Reflexion: Language Agents with Verbal Reinforcement Learning

Reflexion: Language Agents with Verbal Reinforcement Learning

Take the real reward signal from the environment and write it back into the context as words. Reflexion is where a failed episode stops being wasted, which is the seed of everything in the self-improving half of this list.

618Reinforcement Learning
Toolformer: Language Models Can Teach Themselves to Use Tools

Toolformer: Language Models Can Teach Themselves to Use Tools

Where tool calling comes from. Instead of computing five minus three in the weights, the model emits a call and the harness runs it. Declare the tools in the system prompt and the action space is suddenly whatever you are willing to execute.

619Agents
ReAct: Synergizing Reasoning and Acting in Language Models

ReAct: Synergizing Reasoning and Acting in Language Models

Interleave a thought and an action instead of choosing between them. ReAct is the shape almost every agent loop still has, and the talk's point is that models now do this natively, so a modern harness should stop imposing it.

620Agents
WebGPT: Browser-assisted question-answering with human feedback

WebGPT: Browser-assisted question-answering with human feedback

The first time the loop reached outside itself. WebGPT gives a model a browser and human feedback on how it used one, which turns retrieval from a preprocessing step into an action the model chooses to take.

621Agents
Language Models are Unsupervised Multitask Learners

Language Models are Unsupervised Multitask Learners

The V0 harness, and the baseline every later entry is measured against. There is no tool calling here, no skills, no memory: a while-not-EOS loop, top-p sampling, and an environment that scores whatever comes after the delimiter. The talk opens the history here precisely because so little is present, which makes the next six years legible as one move repeated, giving the loop something new it is allowed to do.

622Agents
Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions

Co-evolves environments with the agents that solve them, so the curriculum never settles. This is where open-ended self-improvement starts, years before language models entered the picture.

623Agents
Reinforcing Agents with Collective Skills

Reinforcing Agents with Collective Skills

Binfeng Xu, Yi Dong, Jan Kautz and colleagues at NVIDIA build Skill2Env, a pipeline that compiles public Agent Skills into executable RL environments for terminal agents, each with programmatic tests and a behavioral rubric drawn from the Skill's own quality criteria.

624Agents
182 weeks of AI research · papers per week
Week of Sep 28–Oct 4, 202610 papers →
Apr2023
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2024
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2025
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
Jan2026
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Apr 2023Hover a week to inspect · select to openSep 2026