HyperAIHyperAI

Command Palette

Search for a command to run...

Agent
LLM

Raven: Das Harness der Harnesse für komponierbare agentische Intelligenz

Zusammenfassung

Mit dem Fortschritt großer Sprachmodelle bewegen sich KI-Agenten über isolierte, domänenspezifische Aufgaben hinaus hin zu langfristigen, domänenübergreifenden Arbeitsabläufen. Dieser Übergang offenbart zwei Herausforderungen: Die zunehmende Komplexität von Harnessen macht den manuellen Entwurf schwer skalierbar, während die engere Kopplung an bestimmte Domänen die Allgemeingültigkeit eines einzelnen Harness-Systems einschränkt. Die zentrale Frage verschiebt sich daher von der Entwicklung eines leistungsfähigeren Harness für eine einzelne Domäne hin zur autonomen Konstruktion spezialisierter Harnesse, deren Verbesserung durch Erfahrung und deren domänenübergreifender Orchestrierung. Wir stellen Raven, das Harness der Harnesse, vor – ein quelloffenes Multiagenten-Ökosystem, das modulare Harnesse für bestimmte Modelle und Domänen automatisch konstruiert und weiterentwickelt und dabei jedes ausführbare Modell-Harness-Paar als komponierbare Einheit der Intelligenz behandelt. Zur Unterstützung eines All-Domain-Kollaborationsnetzwerks zerlegt sein Host-Agent Ziele, ordnet Teilaufgaben spezialisierten Agenten zu, koordiniert Ausführungsabhängigkeiten und integriert Ergebnisse, während ein Host-Archiv und EverOS Erfahrungen über Aufgaben hinweg bewahren und Skill Forge diese Erfahrungen als wiederverwendbare Prozeduren verfügbar macht. Unsere Theorie formuliert hinreichende Bedingungen, unter denen eine solche Komposition die zuverlässige Aufgabenabdeckung über die der verfügbaren Einzelagenten hinaus erweitert – bei gemeinsamem Ressourcenbudget. Bei komplexen und langfristigen Aufgaben übertrifft Raven die hochmodernen Agentensysteme deutlich und verschiebt damit die Grenze komponierbarer agentischer Intelligenz.

One-sentence Summary

Researchers at EverMind AI introduce Raven: The Harness of Harnesses, an open-source multi-agent ecosystem that autonomously constructs and evolves modular model-domain harnesses as composable units of intelligence, uses a Host Agent to decompose goals, match subtasks to specialized agents, coordinate dependencies, and integrate results while EverOS and Skill Forge preserve and reuse experience, and provides theoretical sufficient conditions under a shared resource budget for expanded reliable task coverage, significantly outperforming state-of-the-art agent systems on complex long-horizon tasks.

Key Contributions

  • Raven is introduced as an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treats each executable model-harness pair as a composable unit of intelligence, and uses a Host Agent to decompose goals, match subtasks to specialized agents, coordinate execution dependencies, and integrate results across domains.
  • A theoretical framework establishes sufficient conditions under which composing these model-harness units expands reliable task coverage beyond that of the available individual agents under a shared resource budget, accounting for planning and coordination costs.
  • On MAOB, Raven ranks first among compared systems on all four planning metrics under both tested backbones, improving Exact Match over the strongest baseline by 10.4 and 10.5 percentage points; HarnessBank experiments demonstrate benefits from harness adaptation, and SkillCorpus experiments show that a curated skill library improves Raven on three benchmarks, with larger gains than OpenClaw on two of them.

Introduction

The authors address the problem of coordinating LLM-based agents that rely on different tools, execution policies, and specialized harnesses to complete complex multi-stage tasks. An agent’s practical ability depends not only on its model but also on its harness, including tool interfaces, context management, skills, and recovery mechanisms. Prior approaches face scaling challenges in manual harness design and make composition difficult because specialized harnesses encode assumptions about tools, outputs, and handoffs, while multi-agent systems often suffer from inter-agent misalignment and inadequate task verification. The authors introduce Raven, an open-source multi-agent ecosystem that treats each model-harness pair as a unit of composition. A Host Agent matches subtasks to registered native or third-party agents, represents dependencies as a directed acyclic graph, and coordinates execution, while modular harness evolution, persistent memory, and skill reuse support adaptation across tasks. They also provide a theoretical account of when composition expands reliable task coverage and introduce the Multi-Agent Orchestration Benchmark, MAOB, for evaluating orchestration planning.

Dataset

The dataset description covers two components:

SkillCorpus and SkillHub skill catalog

  • Sources: SkillCorpus aggregates public skills through a source registry. SkillHub is the live skill catalog Raven queries. The fixed reference release has 96,401 active skills.
  • Record schema: each skill includes a source-native ID, display name, applicability/trigger description, procedural body, bundled resources, and metadata. In SKILL.md, name and description are frontmatter, the procedure is Markdown, and optional directories hold scripts, references, or assets.
  • Curation pipeline: six stages parse entries, filter malformed or length-ineligible entries, deduplicate, assign quality and issue flags, apply release admission, and attach source priors, embeddings, and index entries.
  • Deduplication: uses content and name fingerprints, then 1,024-dimensional semantic embeddings. Cosine similarity above 0.995 triggers automatic merging; pairs in (0.90, 0.995] receive LLM adjudication.
  • Release and quality: utility, robustness, and safety are scored from 0 to 10 and normalized to 0-1. Hard flags include prompt injection, command injection, unsafe execution, authentication bypass, and CSAM risk. Admission requires a license pass, no malware match, no hard flags, and safety at least 0.3. The 19 issue flags cap quality facets. For admitted skills, content quality is a weighted sum of 0.50 utility, 0.35 robustness, and 0.15 safety, attenuated by safety; structural bonuses add 0.05 for a scripts directory and 0.02 for a references directory. Composite quality combines 0.85 content quality, 0.15 source prior, and the structural bonus, with clipping. Each released skill receives one of sixteen task-domain labels.
  • Retrieval use: the corpus provides a Qwen3-Embedding-0.6B retrieval model and Qwen3-Reranker-0.6B reranker. Retrieval fields are distilled to 3,000 characters from each body. The reference stack recalls and reranks candidates, then an LLM selector returns zero to two skills. Raven combines sources and gives bounded body excerpts to its selection gate.

Multi-Agent Orchestration Benchmark (MAOB)

  • Source and scale: 140 requests modeled on occupational tasks, drawn from 137 distinct occupations and 140 templates. Pattern extraction uses public occupational task collections GDPval and JobBench, but no task text is reused.
  • Composition: each request is paired with a reviewed reference directed acyclic graph over four specialist sub-agents: research, coding, content, and oncall. The benchmark includes all 11 subsets containing at least two domains. Average reference graphs have 2.72 nodes and 1.84 edges, with 257 reference edges total. 98 tasks are purely serial and 42 allow parallel execution. Domain frequencies are content 106, research 103, coding 102, and oncall 70. By subset size, 66 tasks span two domains, 47 span three, and 27 span four.
  • Construction: generation is graph-first. The reference DAG is fixed before request text. Reference graphs are authored with Claude Opus 5; requests are generated from the graph with GLM-5.2. A leakage filter rejects requests that enumerate steps or name domains outright.
  • Quality control: automatic checks reject inconsistent specifications, role-boundary violations, and field inconsistencies. A further pass removes reference nodes not justified by the request, then reconstructs edges from the transitive closure. Expert review records node-set agreement, ordering ambiguity, and attribution of request elements to nodes.
  • Use: MAOB is used to evaluate multi-agent orchestration by measuring agreement between the host's proposed graph and the reviewed reference graph. All requests are self-contained text without attachments.

Method

The authors present Raven, a system that composes executable model-harness pairs through a host agent controlling assignment, information exchange, and execution under a common resource budget.

As shown in the figure below:

The framework formalizes composable agentic intelligence where a model and its harness form a callable execution unit. The host coordinates heterogeneous agents, such as research, code, and design specialists, managing explicit artifact handoffs and control flow.

For task decomposition, the host utilizes staged orchestration guidance. It carries only a summary of the guide initially and loads the full guide only when a request requires multiple agents. The host then composes a typed dependency graph.

Refer to the framework diagram:

The graph planning and admission process ensures structural and capability consistency. A submission includes a flat node list with graph-level flags. The runtime performs five sequential checks covering format, graph structure, agent capability, status, and environment constraints before dispatching any worker. Rejected graphs return the first error with a pointer back to the guide, allowing the host to retry.

Once admitted, the system proceeds to host-mediated graph execution.

As shown in the figure below:

The node lifecycle dictates that a node runs only after all predecessors have completed and settled. An independent LLM call judges the output and transcript tail. A negative verdict suspends the node in an exception state, prompting a host decision to continue, abandon, or replan. Additionally, if a worker raises a clarification request, the host attempts to autofill the answer using the current conversation and memory before deferring to the user.

Artifact and memory flow are managed through file-backed records to avoid context duplication.

Refer to the framework diagram:

A node's prompt is rendered from its template, upstream outputs, memory paths, and files. The runtime writes the prompt, output, and transcript to disk, while memory recording proceeds asynchronously. Completed nodes enter a session-wide node registry, enabling later graphs to reference earlier outputs and single delegations to continue existing agent instances.

To incorporate broader feedback across rounds, the system employs a shared group memory layer.

As shown in the figure below:

The host acts as a sole broker, mediating access to owner-partitioned libraries. It retrieves experience before planning and deposits round summaries. At dispatch, the host prefetches shared experience and recent verdicts for the worker. Assessment occurs in two phases: a close verdict is generated immediately after execution based on an attribution block extracted from the host's reply, and a feedback verdict is created on the next user turn to revise the outcome based on user reaction.

Beyond orchestration, the authors design a modular harness self-evolution mechanism. The policy surrounding a frozen model is exposed through strategy interfaces for memory, planning, capability, and action. An Evolver Agent diagnoses execution failures from evaluation logs and proposes patches, which are screened via paired improvement statistics before being retained in a semantic gene bank. Concurrently, the Skill Forge module retrieves reusable procedures from a curated catalog and local memory, while an experience-driven pipeline extracts execution cases to incrementally cluster and update skills for future tasks.

Experiment

The evaluation examines Raven across orchestration planning, harness self-evolution, four specialist workflows, and procedural skill reuse. Planning is isolated on the Multi-Agent Orchestration Benchmark under matched backbones, where Raven consistently produces better specialist selections and dependency structures than baselines. Harness self-evolution improves held-out performance across seven benchmarks with a frozen model, while the research, coding, design, and oncall specialists each outperform stronger baselines on domain tasks, often with comparable or lower resource use. Skill retrieval experiments further show consistent gains when procedures are added to a fixed harness and backbone, with Raven translating the same retrieved skills into larger improvements than the comparison harness.

Raven outperforms the best same-backbone baseline on all four multi-agent orchestration metrics for both evaluated backbones. DeepSeek-V4-Flash-0731 achieves the strongest absolute Node F1, Edge F1, and Exact Match Rate, while Qwen3.8-27B achieves the strongest Partial Order Accuracy and the largest single-metric gain. Exact Match Rate is the lowest-scoring metric for both backbones and improves by roughly 0.10 in each case. Raven improves all four reported metrics over the best baseline for Qwen3.8-27B and DeepSeek-V4-Flash-0731. DeepSeek-V4-Flash-0731 leads in absolute Node F1, Edge F1, and Exact Match Rate; Qwen3.8-27B leads in Partial Order Accuracy. The largest absolute gain is on Partial Order Accuracy for Qwen3.8-27B, while DeepSeek-V4-Flash-0731 gains most on Node F1. Exact Match Rate remains the weakest absolute metric for both backbones, with nearly equal gains across the two models.

Before a multi-agent graph is dispatched, Raven applies admission checks across format, graph structure, agent capability, agent status, and runtime environment. These checks catch malformed arguments, invalid or cyclic dependencies, missing placeholder backing, unauthorized path use, unavailable agents, and disallowed delegation states. Validation stops at the first failure and returns a pointer to the orchestration guide. Admission checks are grouped into format, graph structure, agent capability, agent status, and environment categories. Graph-structure validation rejects duplicate or previously used identifiers, unresolved dependencies, unbacked output placeholders, invalid path forms, and cycles. Agent capability rules restrict shared instances to stateful agents, limit skill overrides to the head of an instance chain, and block path arguments to agents without local-file access. The first failed check stops validation and directs the caller to the orchestration guide for remediation.

Within the group memory layer, access is centralized in the host, which can read and write its own shared library and every mapped agent library but has no access to the user library. Mapped workers receive read-only access through prefetch to the host library and their own verdicts. Unmapped agents receive no access to these group memory libraries. The host has read/write access to the shared host library and all mapped agent libraries. The host and mapped workers have no access to the user library. Mapped workers receive read-only prefetch access to the host library and to their own verdicts. Unmapped agents have no access to the host library, user library, or agent libraries.

The benchmark covers every possible combination of at least two specialist domains from the four-domain roster, spanning two-domain, three-domain, and four-domain tasks. Most tasks involve two domains, with fewer tasks requiring three domains and the fewest requiring all four. In total, the suite contains 140 tasks across 11 distinct domain combinations. All 11 possible domain subsets with at least two domains are represented. Two-domain tasks are the most common, followed by three-domain and four-domain tasks. The benchmark includes 140 tasks in total.

Raven-Research obtains the highest pooled accuracy among the compared research harnesses on DeepResearch Mixed for all shared backbones. Its lead over the strongest baseline ranges from about three to eight percentage points, with statistical support in most paired comparisons. The advantage is concentrated in BrowseComp and Humanity’s Last Exam, while results on FRAMES are close and DeepSeek-Harness performs better on xBench-DeepSearch. Raven-Research ranks first in pooled accuracy within every shared-backbone group. Paired statistical tests favor Raven-Research in five of six same-backbone comparisons, with the remaining comparison still positive. Benchmark-level gains are concentrated in BrowseComp and Humanity’s Last Exam; on xBench-DeepSearch, DeepSeek-Harness leads slightly.

The experiments evaluate Raven across multi-agent orchestration, admission validation, group memory access control, benchmark coverage, and research-harness accuracy. Raven improves all four orchestration metrics over the best same-backbone baselines, with exact match remaining the weakest metric, while admission checks catch format, graph structure, agent capability, agent status, and environment errors before dispatch. Group memory access is centrally enforced so the host has read/write access to shared and mapped agent libraries, mapped workers receive read-only prefetch access, and unmapped agents have no access. The benchmark covers 140 tasks across all eleven possible combinations of two to four specialist domains, and Raven-Research achieves the highest pooled accuracy within every shared-backbone group, with gains concentrated in BrowseComp and Humanity’s Last Exam.


KI mit KI entwickeln

Von der Idee bis zum Launch – beschleunigen Sie Ihre KI-Entwicklung mit kostenlosem KI-Co-Coding, sofort einsatzbereiter Umgebung und bestem GPU-Preis.

KI-gestütztes kollaboratives Programmieren
Sofort einsatzbereite GPUs
Die besten Preise

HyperAI Newsletters

Abonnieren Sie unsere neuesten Updates
Wir werden die neuesten Updates der Woche in Ihren Posteingang liefern um neun Uhr jeden Montagmorgen
Unterstützt von MailChimp