Command Palette
Search for a command to run...
Raven: 構成可能なエージェント的知能のためのハーネスのハーネス
Raven: 構成可能なエージェント的知能のためのハーネスのハーネス
概要
大規模言語モデルの進歩に伴い、AIエージェントは孤立したドメイン固有のタスクを超え、長期的かつ領域横断的なワークフローへと移行しつつある。この移行により二つの課題が顕在化する。すなわち、ハーネスの複雑化によって手動設計の規模拡大が困難になる一方で、特定ドメインへの密結合が単一ハーネスの汎用性を制限するという課題である。したがって中心的な問いは、一つのドメイン向けにより強力なハーネスをどう設計するかから、専門化されたハーネスをいかに自律的に構築し、経験を通じて改善し、ドメインを横断して編成するかへと移る。我々はRaven(The Harness of Harnesses)を導入する。これはオープンソースのマルチエージェントエコシステムであり、特定のモデルとドメイン向けにモジュール式ハーネスを自動的に構築・進化させ、実行可能なモデル–ハーネスの各ペアを構成可能な知能の単位として扱う。全領域協調ネットワーク(All-Domain Collaboration Network)を支援するため、そのHost Agentは目標を分解し、部分タスクを専門エージェントに割り当て、実行依存関係を調整し、結果を統合する。また、ホストアーカイブとEverOSがタスクをまたいで経験を保持し、Skill Forgeがその経験を再利用可能な手順として利用可能にする。我々の理論は、共有資源予算の下で、このような構成が利用可能な個々のエージェントの信頼できるタスク範囲を超えて拡大するための十分条件を確立する。複雑で長期的なタスクにおいて、Ravenは最先端のエージェントシステムを有意に上回り、構成可能なエージェント的知能の最前線を押し広げる。
One-sentence Summary
Researchers at EverMind AI introduce Raven: The Harness of Harnesses, an open-source multi-agent ecosystem that autonomously constructs and evolves modular model-domain harnesses as composable units of intelligence, uses a Host Agent to decompose goals, match subtasks to specialized agents, coordinate dependencies, and integrate results while EverOS and Skill Forge preserve and reuse experience, and provides theoretical sufficient conditions under a shared resource budget for expanded reliable task coverage, significantly outperforming state-of-the-art agent systems on complex long-horizon tasks.
Key Contributions
- Raven is introduced as an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treats each executable model-harness pair as a composable unit of intelligence, and uses a Host Agent to decompose goals, match subtasks to specialized agents, coordinate execution dependencies, and integrate results across domains.
- A theoretical framework establishes sufficient conditions under which composing these model-harness units expands reliable task coverage beyond that of the available individual agents under a shared resource budget, accounting for planning and coordination costs.
- On MAOB, Raven ranks first among compared systems on all four planning metrics under both tested backbones, improving Exact Match over the strongest baseline by 10.4 and 10.5 percentage points; HarnessBank experiments demonstrate benefits from harness adaptation, and SkillCorpus experiments show that a curated skill library improves Raven on three benchmarks, with larger gains than OpenClaw on two of them.
Introduction
The authors address the problem of coordinating LLM-based agents that rely on different tools, execution policies, and specialized harnesses to complete complex multi-stage tasks. An agent’s practical ability depends not only on its model but also on its harness, including tool interfaces, context management, skills, and recovery mechanisms. Prior approaches face scaling challenges in manual harness design and make composition difficult because specialized harnesses encode assumptions about tools, outputs, and handoffs, while multi-agent systems often suffer from inter-agent misalignment and inadequate task verification. The authors introduce Raven, an open-source multi-agent ecosystem that treats each model-harness pair as a unit of composition. A Host Agent matches subtasks to registered native or third-party agents, represents dependencies as a directed acyclic graph, and coordinates execution, while modular harness evolution, persistent memory, and skill reuse support adaptation across tasks. They also provide a theoretical account of when composition expands reliable task coverage and introduce the Multi-Agent Orchestration Benchmark, MAOB, for evaluating orchestration planning.
Dataset
The dataset description covers two components:
SkillCorpus and SkillHub skill catalog
- Sources: SkillCorpus aggregates public skills through a source registry. SkillHub is the live skill catalog Raven queries. The fixed reference release has 96,401 active skills.
- Record schema: each skill includes a source-native ID, display name, applicability/trigger description, procedural body, bundled resources, and metadata. In SKILL.md, name and description are frontmatter, the procedure is Markdown, and optional directories hold scripts, references, or assets.
- Curation pipeline: six stages parse entries, filter malformed or length-ineligible entries, deduplicate, assign quality and issue flags, apply release admission, and attach source priors, embeddings, and index entries.
- Deduplication: uses content and name fingerprints, then 1,024-dimensional semantic embeddings. Cosine similarity above 0.995 triggers automatic merging; pairs in (0.90, 0.995] receive LLM adjudication.
- Release and quality: utility, robustness, and safety are scored from 0 to 10 and normalized to 0-1. Hard flags include prompt injection, command injection, unsafe execution, authentication bypass, and CSAM risk. Admission requires a license pass, no malware match, no hard flags, and safety at least 0.3. The 19 issue flags cap quality facets. For admitted skills, content quality is a weighted sum of 0.50 utility, 0.35 robustness, and 0.15 safety, attenuated by safety; structural bonuses add 0.05 for a scripts directory and 0.02 for a references directory. Composite quality combines 0.85 content quality, 0.15 source prior, and the structural bonus, with clipping. Each released skill receives one of sixteen task-domain labels.
- Retrieval use: the corpus provides a Qwen3-Embedding-0.6B retrieval model and Qwen3-Reranker-0.6B reranker. Retrieval fields are distilled to 3,000 characters from each body. The reference stack recalls and reranks candidates, then an LLM selector returns zero to two skills. Raven combines sources and gives bounded body excerpts to its selection gate.
Multi-Agent Orchestration Benchmark (MAOB)
- Source and scale: 140 requests modeled on occupational tasks, drawn from 137 distinct occupations and 140 templates. Pattern extraction uses public occupational task collections GDPval and JobBench, but no task text is reused.
- Composition: each request is paired with a reviewed reference directed acyclic graph over four specialist sub-agents: research, coding, content, and oncall. The benchmark includes all 11 subsets containing at least two domains. Average reference graphs have 2.72 nodes and 1.84 edges, with 257 reference edges total. 98 tasks are purely serial and 42 allow parallel execution. Domain frequencies are content 106, research 103, coding 102, and oncall 70. By subset size, 66 tasks span two domains, 47 span three, and 27 span four.
- Construction: generation is graph-first. The reference DAG is fixed before request text. Reference graphs are authored with Claude Opus 5; requests are generated from the graph with GLM-5.2. A leakage filter rejects requests that enumerate steps or name domains outright.
- Quality control: automatic checks reject inconsistent specifications, role-boundary violations, and field inconsistencies. A further pass removes reference nodes not justified by the request, then reconstructs edges from the transitive closure. Expert review records node-set agreement, ordering ambiguity, and attribution of request elements to nodes.
- Use: MAOB is used to evaluate multi-agent orchestration by measuring agreement between the host's proposed graph and the reviewed reference graph. All requests are self-contained text without attachments.
Method
The authors present Raven, a system that composes executable model-harness pairs through a host agent controlling assignment, information exchange, and execution under a common resource budget.
As shown in the figure below:
The framework formalizes composable agentic intelligence where a model and its harness form a callable execution unit. The host coordinates heterogeneous agents, such as research, code, and design specialists, managing explicit artifact handoffs and control flow.
For task decomposition, the host utilizes staged orchestration guidance. It carries only a summary of the guide initially and loads the full guide only when a request requires multiple agents. The host then composes a typed dependency graph.
Refer to the framework diagram:
The graph planning and admission process ensures structural and capability consistency. A submission includes a flat node list with graph-level flags. The runtime performs five sequential checks covering format, graph structure, agent capability, status, and environment constraints before dispatching any worker. Rejected graphs return the first error with a pointer back to the guide, allowing the host to retry.
Once admitted, the system proceeds to host-mediated graph execution.
As shown in the figure below:
The node lifecycle dictates that a node runs only after all predecessors have completed and settled. An independent LLM call judges the output and transcript tail. A negative verdict suspends the node in an exception state, prompting a host decision to continue, abandon, or replan. Additionally, if a worker raises a clarification request, the host attempts to autofill the answer using the current conversation and memory before deferring to the user.
Artifact and memory flow are managed through file-backed records to avoid context duplication.
Refer to the framework diagram:
A node's prompt is rendered from its template, upstream outputs, memory paths, and files. The runtime writes the prompt, output, and transcript to disk, while memory recording proceeds asynchronously. Completed nodes enter a session-wide node registry, enabling later graphs to reference earlier outputs and single delegations to continue existing agent instances.
To incorporate broader feedback across rounds, the system employs a shared group memory layer.
As shown in the figure below:
The host acts as a sole broker, mediating access to owner-partitioned libraries. It retrieves experience before planning and deposits round summaries. At dispatch, the host prefetches shared experience and recent verdicts for the worker. Assessment occurs in two phases: a close verdict is generated immediately after execution based on an attribution block extracted from the host's reply, and a feedback verdict is created on the next user turn to revise the outcome based on user reaction.
Beyond orchestration, the authors design a modular harness self-evolution mechanism. The policy surrounding a frozen model is exposed through strategy interfaces for memory, planning, capability, and action. An Evolver Agent diagnoses execution failures from evaluation logs and proposes patches, which are screened via paired improvement statistics before being retained in a semantic gene bank. Concurrently, the Skill Forge module retrieves reusable procedures from a curated catalog and local memory, while an experience-driven pipeline extracts execution cases to incrementally cluster and update skills for future tasks.
Experiment
The evaluation examines Raven across orchestration planning, harness self-evolution, four specialist workflows, and procedural skill reuse. Planning is isolated on the Multi-Agent Orchestration Benchmark under matched backbones, where Raven consistently produces better specialist selections and dependency structures than baselines. Harness self-evolution improves held-out performance across seven benchmarks with a frozen model, while the research, coding, design, and oncall specialists each outperform stronger baselines on domain tasks, often with comparable or lower resource use. Skill retrieval experiments further show consistent gains when procedures are added to a fixed harness and backbone, with Raven translating the same retrieved skills into larger improvements than the comparison harness.
Raven outperforms the best same-backbone baseline on all four multi-agent orchestration metrics for both evaluated backbones. DeepSeek-V4-Flash-0731 achieves the strongest absolute Node F1, Edge F1, and Exact Match Rate, while Qwen3.8-27B achieves the strongest Partial Order Accuracy and the largest single-metric gain. Exact Match Rate is the lowest-scoring metric for both backbones and improves by roughly 0.10 in each case. Raven improves all four reported metrics over the best baseline for Qwen3.8-27B and DeepSeek-V4-Flash-0731. DeepSeek-V4-Flash-0731 leads in absolute Node F1, Edge F1, and Exact Match Rate; Qwen3.8-27B leads in Partial Order Accuracy. The largest absolute gain is on Partial Order Accuracy for Qwen3.8-27B, while DeepSeek-V4-Flash-0731 gains most on Node F1. Exact Match Rate remains the weakest absolute metric for both backbones, with nearly equal gains across the two models.
Before a multi-agent graph is dispatched, Raven applies admission checks across format, graph structure, agent capability, agent status, and runtime environment. These checks catch malformed arguments, invalid or cyclic dependencies, missing placeholder backing, unauthorized path use, unavailable agents, and disallowed delegation states. Validation stops at the first failure and returns a pointer to the orchestration guide. Admission checks are grouped into format, graph structure, agent capability, agent status, and environment categories. Graph-structure validation rejects duplicate or previously used identifiers, unresolved dependencies, unbacked output placeholders, invalid path forms, and cycles. Agent capability rules restrict shared instances to stateful agents, limit skill overrides to the head of an instance chain, and block path arguments to agents without local-file access. The first failed check stops validation and directs the caller to the orchestration guide for remediation.
Within the group memory layer, access is centralized in the host, which can read and write its own shared library and every mapped agent library but has no access to the user library. Mapped workers receive read-only access through prefetch to the host library and their own verdicts. Unmapped agents receive no access to these group memory libraries. The host has read/write access to the shared host library and all mapped agent libraries. The host and mapped workers have no access to the user library. Mapped workers receive read-only prefetch access to the host library and to their own verdicts. Unmapped agents have no access to the host library, user library, or agent libraries.
The benchmark covers every possible combination of at least two specialist domains from the four-domain roster, spanning two-domain, three-domain, and four-domain tasks. Most tasks involve two domains, with fewer tasks requiring three domains and the fewest requiring all four. In total, the suite contains 140 tasks across 11 distinct domain combinations. All 11 possible domain subsets with at least two domains are represented. Two-domain tasks are the most common, followed by three-domain and four-domain tasks. The benchmark includes 140 tasks in total.
Raven-Research obtains the highest pooled accuracy among the compared research harnesses on DeepResearch Mixed for all shared backbones. Its lead over the strongest baseline ranges from about three to eight percentage points, with statistical support in most paired comparisons. The advantage is concentrated in BrowseComp and Humanity’s Last Exam, while results on FRAMES are close and DeepSeek-Harness performs better on xBench-DeepSearch. Raven-Research ranks first in pooled accuracy within every shared-backbone group. Paired statistical tests favor Raven-Research in five of six same-backbone comparisons, with the remaining comparison still positive. Benchmark-level gains are concentrated in BrowseComp and Humanity’s Last Exam; on xBench-DeepSearch, DeepSeek-Harness leads slightly.
The experiments evaluate Raven across multi-agent orchestration, admission validation, group memory access control, benchmark coverage, and research-harness accuracy. Raven improves all four orchestration metrics over the best same-backbone baselines, with exact match remaining the weakest metric, while admission checks catch format, graph structure, agent capability, agent status, and environment errors before dispatch. Group memory access is centrally enforced so the host has read/write access to shared and mapped agent libraries, mapped workers receive read-only prefetch access, and unmapped agents have no access. The benchmark covers 140 tasks across all eleven possible combinations of two to four specialist domains, and Raven-Research achieves the highest pooled accuracy within every shared-backbone group, with gains concentrated in BrowseComp and Humanity’s Last Exam.