HyperAIHyperAI

Command Palette

Search for a command to run...

ASI-Bench:人工超知能の黎明において

概要

人工超知能(ASI)は、AIが既存知識の習得を超えて、未知を探求し、新たな知識を創造し、新しいアイデアを検証可能な成果へと転換することを必要とする。しかし、今日のAIシステムの能力は、依然として既存の人間の知識の学習・圧縮・適用に大きく依拠している。それに応じて、既存のベンチマークは主に、AIが学習した知識に基づいて正しい答えを出せるか、あるいは人間の広範な指導の下でタスクを遂行できるかをテストしている。そこで我々は、一般的な研究領域にわたってAIシステムの革新的探索能力と自律的科学実行能力を統合的に評価する初のベンチマークであり、同一研究プロジェクト内で人間による方法論的指導を段階的に撤去し、AIがどこまで自律的に進めるかをテストする初のベンチマークであるASI-Benchを導入する。40名を超える専門家によって31,000人時以上のコストをかけて構築されたASI-Benchは、11の科学領域にわたる60のプロジェクトレベルの研究タスクを含み、AIが独立して手法を選択し、研究を遂行し、検証可能な成果を生み出せるかをテストするために方法論的指導を段階的に削減する。すべてのタスクは専門家レビュー、AI支援監査、サンドボックス実行、採点者検証を経ている。18の最先端エージェント・モデル構成において、平均スコアは完全な方法論的指導がある場合の50.91から、手法のみが指定された場合の29.10、エージェント自身が手法を決定しなければならない場合の26.62へと低下する。この急激な低下は、現在のシステムが依然として人間の指導に大きく依存しており、エンドツーエンドのプロジェクトレベルの科学研究を自律的に遂行するには程遠いことを示している。ASI-Benchは世界に開かれている。我々は、あらゆる場所の研究者および開発者に対し、新しいタスクの提供、今日のAIの限界への挑戦、そしてhttps://asibench.apexin.ai/submit における人工超知能への人類の集団的な道のりを加速することへの貢献を呼びかける。

One-sentence Summary

Researchers from Tsinghua University, Massachusetts Institute of Technology, Harvard University, et al. introduce ASI-Bench, the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution across 60 project-level research tasks in 11 scientific disciplines by progressively withdrawing methodological guidance, with 18 agent-model configurations dropping from 50.91 under full guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves.

Key Contributions

  • Introduces ASI-Bench, the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution across general research domains, containing 60 project-level research tasks across 11 scientific domains and built with over 40 experts and 31,000+ human hours.
  • Adds a progressive guidance-withdrawal design that reduces human methodological guidance within the same research project to test whether AI systems can independently select methods, conduct research, and produce verifiable results, with validation through expert review, AI-assisted auditing, sandbox execution, and scorer validation.
  • Reports evaluation across 18 state-of-the-art agent-model configurations, with average performance dropping from 50.91 with full methodological guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves, showing current systems remain heavily dependent on human guidance.

Introduction

Toward artificial superintelligence, AI systems must move beyond applying existing human knowledge and begin exploring open-ended scientific problems with verifiable results. Prior benchmarks tend to evaluate known-answer tasks, human-specified procedures, or isolated research components, so they offer limited evidence about autonomous end-to-end discovery. The authors introduce ASI-Bench, a benchmark of 60 project-level scientific tasks across 11 domains with executable environments and verifiable research artifacts. Its B1 to B4 guidance gradient progressively withdraws methodological support, and evaluation of 18 state-of-the-art agent and model configurations shows average performance falling from 50.91 with full guidance to 26.62 when agents must determine the method themselves, revealing a substantial gap between scientific execution and autonomous research.

Dataset

ASI-Bench Dataset Description

The authors present ASI-Bench as an evaluation benchmark rather than a training set. It is built to test end-to-end scientific research capability under progressively reduced human methodological guidance.

Composition and sources

  • ASI-Bench contains 60 project-level research tasks.
  • The tasks span 11 scientific domains: mathematics, physics, chemistry, biology, astronomy, materials science, earth science, medicine and biostatistics, computer science, robotics, and electrical engineering.
  • Candidate tasks come from more than 1,300 research ideas collected from scientific sources.
  • The provided sections do not report per-domain task counts.

Guidance variants and task schema

  • Each task is presented under matched guidance conditions while keeping the underlying task, data, required outputs, and scoring criteria fixed.
  • B1 gives the full governing equations, numerical formulation, and solver procedure.
  • B2 removes the full procedure but gives methodological guidance about the problem class and suitable numerical approaches.
  • B3 removes methodological guidance and gives only observed data, the scientific objective, and the required outputs.
  • B4 keeps the B3 setting but adds plausible task-irrelevant information to test whether the agent can ignore distraction.
  • Each task involves problem understanding, method selection, implementation, experimentation, failure diagnosis, iterative refinement, and result validation.

Processing and filtering

  • Construction begins with more than 1,300 candidate research ideas.
  • The process includes five review rounds, over 1,100 review assignments, and more than 2,000 task revisions.
  • Reviewers examine scientific formulation, task specification, B1-B4 information design, reference results, evaluation criteria, information leakage, and unintended shortcuts.
  • More than 31,000 human-hours were invested in construction and validation.
  • Retained tasks are validated through over 1,500 sandbox runs for runtime stability, reference reproducibility, artifact generation, and scoring consistency.
  • Tasks with unresolved scientific errors, unstable execution, evaluation misalignment, or unintended solution paths are revised or excluded.

Usage

  • ASI-Bench is used for evaluation, not as a training corpus.
  • No training split or mixture ratios are described.
  • Systems are evaluated under the same underlying task, data, required outputs, and scoring criteria, while only the guidance level changes.
  • The benchmark can compare different backbone models within the same agent framework or different agents with the same model.
  • Across 18 state-of-the-art Agent x Model configurations, the average B3 score is 26.62.

Cropping and metadata

  • No image or text cropping strategy is described in the provided sections.
  • Tasks use task-specific datasets and project-level inputs rather than uniform crops.
  • The main metadata construction is the B1-B4 guidance design, required outputs, and scoring criteria.

Method

The authors design ASI-Bench to evaluate the extent to which AI can conduct scientific research as human methodological guidance is progressively withdrawn. The benchmark consists of 60 project-level research tasks spanning 11 scientific domains. Each project is evaluated under matched guidance conditions while keeping the underlying task, data, required outputs, and scoring criteria fixed.

The construction of ASI-Bench follows a rigorous, large-scale iterative pipeline rather than a single-pass collection process. As shown in the framework diagram below:

The process begins with problem collection from various scientific sources, yielding over 1,300 candidate research ideas. These candidates undergo extensive expert reviews, comprising five review rounds, more than 1,100 review assignments, and over 2,000 task revisions. Reviewers examine the scientific formulation, task specification, information design, reference results, and evaluation criteria. They also check for information leakage and unintended shortcuts that could lead to high scores without correctly solving the task. This construction and validation process required more than 31,000 human-hours.

Following the review process, each retained task is validated through end-to-end execution in isolated sandboxes. The authors conduct more than 1,500 sandbox runs to verify runtime stability, reference reproducibility, artifact generation, and scoring consistency. Tasks with unresolved scientific errors, unstable execution, evaluation misalignment, or unintended solution paths are revised or excluded, resulting in the final benchmark of 60 project-level tasks.

To measure scientific autonomy, the benchmark progressively reduces human methodological guidance. Each task is designed as a complex, project-level scientific investigation rather than an isolated question. Starting from a research objective and task-specific data, agents must carry out a long-horizon, multi-stage research process spanning problem understanding, method selection, implementation, experimentation, failure diagnosis, iterative refinement, and result validation. Across the 60 tasks, completing these research processes involves more than 2,600 interaction turns and 2,400 execution steps.

The guidance progression shifts scientific responsibility from humans to AI. In the most guided condition, the governing equations, numerical formulation, and solver procedure are explicitly provided, requiring the agent to mainly implement and execute the prescribed approach. In subsequent conditions, the full procedure is removed, leaving only methodological guidance about the class of problem and suitable approaches. In the least guided condition, the agent receives only the observed data, the scientific objective, and the required outputs, forcing it to determine the underlying model, choose an appropriate numerical method, implement it, and validate the resulting prediction. A final condition retains this minimal guidance but adds plausible yet task-irrelevant information to test whether the agent can maintain its research direction under distraction.

The benchmark covers fundamental science, life science, computing, and engineering, requiring the same agent and model system to generalize across different data types, scientific methods, and validation criteria. This breadth tests whether autonomous research transfers across disciplines rather than remaining limited to a single domain.

Experiment

ASI-Bench evaluates 18 representative Agent×Model configurations on 60 project-level research tasks across 11 scientific domains under four guidance conditions (B1–B4), where human methodological guidance is progressively withdrawn. The main results show that current systems remain far from reliable autonomous scientific discovery, with only the strongest configuration achieving a B3 score above 50, and that the sharp drop from B1 to B2 indicates the primary bottleneck is turning a selected method into a complete research procedure rather than method choice or distraction. Results also show that scientific capability emerges from the interaction between the backbone model and the agent harness. Computational-cost experiments further find that complete guidance reduces token and time costs while incomplete methodological guidance can increase overhead, and that higher spending does not reliably translate into better scientific performance.

Representative benchmarks for advanced AI capability each emphasize different aspects of autonomous research, such as broad academic knowledge, scientific coding, tool-based terminal tasks, or research replication. Cross-domain generality, method autonomy, and end-to-end research are rarely combined in a single benchmark, and none of the compared benchmarks explicitly evaluates a guidance gradient across progressively reduced methodological support. This leaves a gap for joint evaluation of general intelligence, independent method selection, and autonomous execution. Most existing benchmarks cover only one or two dimensions of autonomous research, with explicit cross-domain generality limited to a few settings. End-to-end research coverage appears in benchmarks focused on research replication, ML engineering, and AI R&D, while method autonomy is partial or absent in most others. No representative benchmark in the comparison explicitly evaluates a guidance gradient that varies the level of human methodological guidance.

Performance drops most sharply when detailed procedural guidance is removed, while removing the method choice causes a much smaller additional decline and irrelevant context has little effect. Even the strongest configuration reaches only moderate autonomous performance, though stronger inference-time reasoning offers a notable gain. Harness choice can substantially change the same model's capability, but the effect varies across model-harness pairs. The sharpest average decline comes from losing step-by-step procedural guidance, not from losing the method choice. Stronger inference-time reasoning improves autonomous method selection, but the best configuration remains moderate, and harness choices can markedly shift model scores.

The first analysis compares representative autonomous research benchmarks and finds that they rarely combine cross-domain generality, method autonomy, and end-to-end research, while none explicitly evaluates a guidance gradient with varying methodological support. The second experiment measures model performance as guidance is reduced, showing the largest drop when detailed procedural guidance is removed, a smaller additional decline when method choice is removed, and little effect from irrelevant context. Stronger inference-time reasoning improves autonomous method selection but still leaves the best configuration at moderate performance, and harness choice can substantially shift model capability.


AIでAIを構築

アイデアからローンチまで — 無料のAIコーディング支援、すぐに使える環境、最高のGPU価格でAI開発を加速。

AI コーディング補助
すぐに使える GPU
最適な料金体系

HyperAI Newsletters

最新情報を購読する
北京時間 毎週月曜日の午前9時 に、その週の最新情報をメールでお届けします
メール配信サービスは MailChimp によって提供されています