Papers
arxiv:2609.33439

Raven: The Harness of Harnesses for Composable Agentic Intelligence

Published on Sep 27
· Submitted by
LivXue
on Sep 30
#1 Paper of the day
Authors:

Abstract

As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, The Harness of Harnesses, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an All-Domain Collaboration Network, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.

Community

Paper submitter

As large language models advance, AI agents are moving beyond isolated, domain-specific tasks toward long-horizon, cross-domain workflows. This transition exposes two challenges: increasing harness complexity makes manual design difficult to scale, while tighter coupling to specific domains limits the generality of a single harness. The central question thus shifts from how to engineer a stronger harness for one domain to how to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains. We introduce Raven, \emph{The Harness of Harnesses}, an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains, treating each executable model--harness pair as a composable unit of intelligence. To support an \emph{All-Domain Collaboration Network}, its Host Agent decomposes goals, matches subtasks to specialized agents, coordinates execution dependencies, and integrates results, while a host archive and EverOS preserve experience across tasks and Skill Forge makes that experience available as reusable procedures. Our theory establishes sufficient conditions for such composition to expand reliable task coverage beyond that of the available individual agents under a shared resource budget. On complex and long-horizon tasks, Raven significantly outperforms the state-of-the-art agent systems, pushing the frontier of composable agentic intelligence.

Good job, innovative ideas.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Congrats on the release. I've only read the abstract, so forgive me if the paper answers this. When an experience in the archive gets turned into a reusable procedure in Skill Forge, what has to be true before it's promoted? I'm curious whether the archive records who or what judged the lesson was right, or only that the run succeeded. A run can succeed for the wrong reason, and then the wrong reason gets reused.

·

TL;DR: Skill Forge combines LLM review of experience with ongoing skill revision, confidence updates, and a designed policy for retiring low-confidence skills.
Before an experience becomes a skill, the LLM examines the trajectory to identify reusable reasoning, consequential decisions, and verification steps. The prompts ask why an approach worked and discourage procedures that only fit one particular case. The model can create a skill, update an existing one, or make no change.
Each skill retains confidence and links to its supporting cases. Later related cases are considered alongside the existing skill and its historical evidence. Confirming cases can raise confidence; contradictions can lower it, prompt revisions, or add specific pitfalls. Failed cases also contribute corrective lessons. The lifecycle design additionally includes maturity assessment and retirement when confidence falls below 0.1.
The judgment comes from the LLM extraction process, and the archive preserves the approach, extracted insight, quality estimate, and supporting-case lineage. These mechanisms help limit the reuse of misleading lessons and allow skills to improve as evidence accumulates. A successful run remains evidence to interpret, rather than sufficient grounds to treat every step as a reliable procedure.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33439
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33439 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33439 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 4