🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 4, 2026
Agents

Raven: The Harness of Harnesses for Composable Agentic Intelligence

First page
Raven: The Harness of Harnesses for Composable Agentic Intelligence
The curator’s take

EverMind AI releases Raven, an open-source multi-agent system that builds and evolves a separate harness for each model and domain, treats every model-harness pair as a composable unit, and coordinates them through a host agent with shared memory and a curated skill library.

Ask this paper

Key points
01

Architecture. A Host Agent decomposes a goal into a graph of subtasks and routes each to a specialized model-harness pair (research, code, design, on-call), with shared group memory passed between them.

02

Harness evolution. Harnesses are modular. Failures are diagnosed, harness components are mutated and recombined from a gene bank, and candidates pass a statistical screening gate before replacing the current version.

03

Raven-Research. With DeepSeek-V4-Flash it scores 69.3% on BrowseComp and 60.0% on Humanity's Last Exam, against at most 62.4% and 43.3% for two other harnesses on the same model, and above three commercial research services (50.3% to 63.2%).

04

Raven-Code. On SWE-Refactor it beats compared leaderboard entries by 9.5 points with DeepSeek-V4-Flash. On SWE-bench Verified the margins are small: 0.6 and 1.2 points, or 3 and 6 extra tasks out of 500.

05

Skill retrieval. The curated skill library with fine-tuned retrieval raises SkillsBench Pass@1 from 9.2% with no skills to 22.6%; swapping in off-the-shelf retrieval drops it to 13.8% and using the raw crawl of about 821,000 skills drops it to 14.9%.

Abstract

As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.

Every Monday
Get next week’s papers.
Subscribe on Substack