AI Papers of the Week
Every paper worth reading in AI, hand-picked one week at a time.
Discover and explore top AI papers with Claude Code or Codex
npx @dair-ai/mcp setup
Paritok-4B: Intent-Conditioned Context Compression for Coding Agents
Jiayu Shi and Luzhuo Chen release Paritok-4B, a 4B LoRA compressor that shrinks coding-agent context to a quarter of its size by extracting spans rather than paraphrasing them, with weights and data open under Apache 2.0.

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents
Shared skill libraries are usually treated as a safe way for coding agents to reuse each other's work, and EvoMal shows they propagate malware. A planted malicious skill is never invoked, but the agent retrieves it as an authoring template, writes a new skill that preserves the payload, and each authored copy re-enters the library to be imitated again. Across six models the self-poisoning rate runs 20.3% to 41.8%, deleting every planted skill does not clean it up, and a counter-prompt discouraging banner-style copying drops it to 6.7%.

Tunable Tool-Call Rates in LLM Agents via Representation Steering
Yuqi Chen, Vincent Siu, Dawn Song and Chenguang Wang show that whether an instruction-tuned model calls a tool is controlled by a single linear direction in its residual stream, extractable without training and steerable at inference.

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses
Recuris splits agent memory in two, with a Working Memory tracking task progress and an Experiential Memory holding skills, so skill selection is grounded in the current task state rather than the full growing history. Because skill use is anchored to an explicit state, a failed run points at a specific memory component, and a fixed Meta-Agent turns that evidence into validation-gated updates to Skill Memory. It improves task success in 35 of 37 completed model-benchmark pairs, adding 17.8 points to GPT-5.6 Sol on tau-bench and taking Claude Opus 5 to 87.9%.

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
FM-Bench turns football club management into a 20-year test of sustained agent decision-making. Fifteen frontier models operate through 26 tools and hundreds of consequential decisions in a deterministic environment with no LLM judge. The results show that model scale, price, vendor, and token spend do not predict performance; managerial behavior and memory discipline do. Every model also fails to learn hidden market prices from repeated feedback, exposing a concrete limit in long-horizon adaptation.

OpenJarvis: Personal AI, On Personal Devices
The same argument pointed at your laptop. Decompose the personal AI stack into five primitives, then let a frontier cloud model search over that spec while everything runs locally at inference. Roughly 800x lower marginal cost, with the harness rather than the model closing the gap.

LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
Long-term memory is crucial for agents in specialized web environments, where success depends on recalling interface affordances, state dynamics, workflows, and recurring failure modes. However, existing memory benchmarks for agents mostly focus on user histories, short traces, or downstream task success, leaving open how to directly evaluate whether memory systems effectively internalize environment-specific experience. To address this gap, we introduce LongMemEval-V2 (LME-V2), a benchmark for evaluating whether memory systems can help agents acquire the experience needed to become knowledgeable colleagues in customized environments. LME-V2 contains 451 manually curated questions covering five core memory abilities for web agents: static state recall, dynamic state tracking, workflow knowledge, environment gotchas, and premise awareness. Questions are paired with history trajectories containing up to 500 trajectories and 115M tokens. We use a context gathering formulation: memory systems consume history trajectories and return compact evidence for downstream question answering. We propose a suite of two memory methods: AgentRunbook-R, an efficient RAG-based memory with knowledge pools for raw state observations, events, and strategy notes, and AgentRunbook-C, which stores trajectories as files and invokes a coding agent to gather evidence in an augmented sandbox. Experiments show that AgentRunbook-C achieves the best performance with 72.5% average accuracy, outperforming the strongest RAG baseline (48.5%) and the off-the-shelf coding agent baseline (69.3%). Despite the strong performance gains, coding agent based methods have high latency costs. While AgentRunbook-C advances the accuracy-latency Pareto frontier, substantial room for improvement remains. Together, these results establish LME-V2 as a challenging testbed for developing long-term memory systems for environment experience.

Continual Harness: Online Adaptation for Self-Improving Foundation Agents
The last step before online learning. Continual Harness keeps history, memory, skills, prompts, and sub-agent specs across trajectories and mutates them while the agent runs, then goes further and updates the weights DAgger-style from what just happened. The presenter calls test-time training the direction that matters most.

PostTrainBench: Can LLM Agents Automate LLM Post-Training?
Hands an agent one base model, one GPU and ten hours, then asks it to post-train the model on its own. It measures AI automating AI training more directly than any other benchmark here, and audits the reward hacking that follows.

Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
Evolves a group of agents that share experience across branches instead of a tree of isolated variants, and reports the strongest coding results in this lineage. The most convincing successor to the Darwin Gödel Machine so far.

Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine
Shows that an agent's benchmark score predicts its descendants' improvement poorly, names that the Metaproductivity-Performance Mismatch, and offers a measurable substitute aggregated over a clade. It answers which variant a self-modification search should expand next.

Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agents
Self-evolution drifts into unsafe behavior along four paths, which are the model, its memory, its tools and its workflow. The first broad safety study of self-evolving agents, and the counterweight to everything above it.

AlphaEvolve: A coding agent for scientific and algorithmic discovery
An evolutionary coding agent that discovered new algorithms and optimized Google's own infrastructure. It accelerated training of the model that powers it, which makes it the clearest published case of a deployed system improving its own production pipeline.

A Self-Improving Coding Agent
A coding agent with ordinary tools edits its own codebase and improves from 17% to 53% on a random subset of SWE-bench Verified, with no gradient updates. The simplest demonstration that one agent can be both the system under repair and the engineer doing it.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Open-ended machine learning research environments scored for both AI agents and human experts under matched time budgets. It is the measurement behind the question the labs are actually asking, which is whether AI can do AI research.

Gödel Agent: A Self-Referential Agent Framework for Recursive Self-Improvement
An agent that reads and rewrites its own logic at runtime, including the logic it uses to rewrite itself, guided only by a high-level objective. It is the first language-model framework to implement the Gödel machine structure directly, with evaluation in place of proof.

Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
Spawning another agent becomes just another tool call. The framing here, agents with roles that persist and can be addressed, is what makes sub-agents an addressable resource rather than a one-shot fan-out.

Reflexion: Language Agents with Verbal Reinforcement Learning
Take the real reward signal from the environment and write it back into the context as words. Reflexion is where a failed episode stops being wasted, which is the seed of everything in the self-improving half of this list.

Toolformer: Language Models Can Teach Themselves to Use Tools
Where tool calling comes from. Instead of computing five minus three in the weights, the model emits a call and the harness runs it. Declare the tools in the system prompt and the action space is suddenly whatever you are willing to execute.

ReAct: Synergizing Reasoning and Acting in Language Models
Interleave a thought and an action instead of choosing between them. ReAct is the shape almost every agent loop still has, and the talk's point is that models now do this natively, so a modern harness should stop imposing it.

WebGPT: Browser-assisted question-answering with human feedback
The first time the loop reached outside itself. WebGPT gives a model a browser and human feedback on how it used one, which turns retrieval from a preprocessing step into an action the model chooses to take.

Language Models are Unsupervised Multitask Learners
The V0 harness, and the baseline every later entry is measured against. There is no tool calling here, no skills, no memory: a while-not-EOS loop, top-p sampling, and an environment that scores whatever comes after the delimiter. The talk opens the history here precisely because so little is present, which makes the next six years legible as one move repeated, giving the loop something new it is allowed to do.

Paired Open-Ended Trailblazer (POET): Endlessly Generating Increasingly Complex and Diverse Learning Environments and Their Solutions
Co-evolves environments with the agents that solve them, so the curriculum never settles. This is where open-ended self-improvement starts, years before language models entered the picture.

Reinforcing Agents with Collective Skills
Binfeng Xu, Yi Dong, Jan Kautz and colleagues at NVIDIA build Skill2Env, a pipeline that compiles public Agent Skills into executable RL environments for terminal agents, each with programmatic tests and a behavioral rubric drawn from the Skill's own quality criteria.