Command Palette
Search for a command to run...
ASI-Bench: 인공 초지능의 여명기에
ASI-Bench: 인공 초지능의 여명기에
초록
인공 초지능(ASI)은 AI가 기존 지식의 습득을 넘어 미지의 영역을 탐구하고, 새로운 지식을 창출하며, 새로운 아이디어를 검증 가능한 결과로 전환할 것을 요구한다. 그러나 오늘날 AI 시스템의 역량은 여전히 기존 인간 지식의 학습, 압축, 적용에 크게 의존하고 있다. 이에 따라 기존 벤치마크는 주로 AI가 학습된 지식을 바탕으로 정답을 도출할 수 있는지, 또는 광범위한 인간의 지도 아래 과제를 완수할 수 있는지를 평가하는 데 그친다. 이에 우리는 일반 연구 분야 전반에 걸쳐 AI 시스템의 혁신적 탐구 능력과 자율적 과학 수행 능력을 통합적으로 평가하는 최초의 벤치마크인 ASI-Bench를 소개한다. 또한 동일한 연구 프로젝트 내에서 인간의 방법론적 지도를 점진적으로 철회함으로써 AI가 어디까지 독자적으로 진행할 수 있는지를 시험하는 최초의 벤치마크이기도 하다. 40명 이상의 전문가가 31,000시간 이상의 인력을 투입하여 구축한 ASI-Bench는 11개 과학 분야에 걸친 60개의 프로젝트 수준 연구 과제를 포함하며, 방법론적 지도를 단계적으로 축소하여 AI가 독립적으로 방법을 선택하고, 연구를 수행하며, 검증 가능한 결과를 산출할 수 있는지를 평가한다. 모든 과제는 전문가 검토, AI 보조 감사, 샌드박스 실행, 채점자 검증을 거친다. 18개의 최신 에이전트-모델 구성에서 평균 점수는 완전한 방법론적 지도가 제공될 때 50.91점에서, 방법만 지정된 경우 29.10점, 에이전트가 스스로 방법을 결정해야 하는 경우 26.62점으로 하락한다. 이러한 급격한 하락은 현재 시스템이 여전히 인간의 지도에 크게 의존하고 있으며, 프로젝트 수준의 과학 연구를 종단 간으로 자율 수행하기에는 아직 크게 미치지 못함을 보여준다. ASI-Bench는 전 세계에 공개되어 있다. 우리는 모든 곳의 연구자와 개발자가 새로운 과제를 기여하고, 오늘날 AI의 한계에 도전하며, 인공 초지능을 향한 인류 공동의 여정을 가속화하는 데 동참하기를 초대한다: https://asibench.apexin.ai/submit.
One-sentence Summary
Researchers from Tsinghua University, Massachusetts Institute of Technology, Harvard University, et al. introduce ASI-Bench, the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution across 60 project-level research tasks in 11 scientific disciplines by progressively withdrawing methodological guidance, with 18 agent-model configurations dropping from 50.91 under full guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves.
Key Contributions
- Introduces ASI-Bench, the first benchmark to jointly evaluate innovative exploration and autonomous scientific execution across general research domains, containing 60 project-level research tasks across 11 scientific domains and built with over 40 experts and 31,000+ human hours.
- Adds a progressive guidance-withdrawal design that reduces human methodological guidance within the same research project to test whether AI systems can independently select methods, conduct research, and produce verifiable results, with validation through expert review, AI-assisted auditing, sandbox execution, and scorer validation.
- Reports evaluation across 18 state-of-the-art agent-model configurations, with average performance dropping from 50.91 with full methodological guidance to 29.10 when only the method is specified and 26.62 when agents must determine the method themselves, showing current systems remain heavily dependent on human guidance.
Introduction
Toward artificial superintelligence, AI systems must move beyond applying existing human knowledge and begin exploring open-ended scientific problems with verifiable results. Prior benchmarks tend to evaluate known-answer tasks, human-specified procedures, or isolated research components, so they offer limited evidence about autonomous end-to-end discovery. The authors introduce ASI-Bench, a benchmark of 60 project-level scientific tasks across 11 domains with executable environments and verifiable research artifacts. Its B1 to B4 guidance gradient progressively withdraws methodological support, and evaluation of 18 state-of-the-art agent and model configurations shows average performance falling from 50.91 with full guidance to 26.62 when agents must determine the method themselves, revealing a substantial gap between scientific execution and autonomous research.
Dataset
ASI-Bench Dataset Description
The authors present ASI-Bench as an evaluation benchmark rather than a training set. It is built to test end-to-end scientific research capability under progressively reduced human methodological guidance.
Composition and sources
- ASI-Bench contains 60 project-level research tasks.
- The tasks span 11 scientific domains: mathematics, physics, chemistry, biology, astronomy, materials science, earth science, medicine and biostatistics, computer science, robotics, and electrical engineering.
- Candidate tasks come from more than 1,300 research ideas collected from scientific sources.
- The provided sections do not report per-domain task counts.
Guidance variants and task schema
- Each task is presented under matched guidance conditions while keeping the underlying task, data, required outputs, and scoring criteria fixed.
- B1 gives the full governing equations, numerical formulation, and solver procedure.
- B2 removes the full procedure but gives methodological guidance about the problem class and suitable numerical approaches.
- B3 removes methodological guidance and gives only observed data, the scientific objective, and the required outputs.
- B4 keeps the B3 setting but adds plausible task-irrelevant information to test whether the agent can ignore distraction.
- Each task involves problem understanding, method selection, implementation, experimentation, failure diagnosis, iterative refinement, and result validation.
Processing and filtering
- Construction begins with more than 1,300 candidate research ideas.
- The process includes five review rounds, over 1,100 review assignments, and more than 2,000 task revisions.
- Reviewers examine scientific formulation, task specification, B1-B4 information design, reference results, evaluation criteria, information leakage, and unintended shortcuts.
- More than 31,000 human-hours were invested in construction and validation.
- Retained tasks are validated through over 1,500 sandbox runs for runtime stability, reference reproducibility, artifact generation, and scoring consistency.
- Tasks with unresolved scientific errors, unstable execution, evaluation misalignment, or unintended solution paths are revised or excluded.
Usage
- ASI-Bench is used for evaluation, not as a training corpus.
- No training split or mixture ratios are described.
- Systems are evaluated under the same underlying task, data, required outputs, and scoring criteria, while only the guidance level changes.
- The benchmark can compare different backbone models within the same agent framework or different agents with the same model.
- Across 18 state-of-the-art Agent x Model configurations, the average B3 score is 26.62.
Cropping and metadata
- No image or text cropping strategy is described in the provided sections.
- Tasks use task-specific datasets and project-level inputs rather than uniform crops.
- The main metadata construction is the B1-B4 guidance design, required outputs, and scoring criteria.
Method
The authors design ASI-Bench to evaluate the extent to which AI can conduct scientific research as human methodological guidance is progressively withdrawn. The benchmark consists of 60 project-level research tasks spanning 11 scientific domains. Each project is evaluated under matched guidance conditions while keeping the underlying task, data, required outputs, and scoring criteria fixed.
The construction of ASI-Bench follows a rigorous, large-scale iterative pipeline rather than a single-pass collection process. As shown in the framework diagram below:
The process begins with problem collection from various scientific sources, yielding over 1,300 candidate research ideas. These candidates undergo extensive expert reviews, comprising five review rounds, more than 1,100 review assignments, and over 2,000 task revisions. Reviewers examine the scientific formulation, task specification, information design, reference results, and evaluation criteria. They also check for information leakage and unintended shortcuts that could lead to high scores without correctly solving the task. This construction and validation process required more than 31,000 human-hours.
Following the review process, each retained task is validated through end-to-end execution in isolated sandboxes. The authors conduct more than 1,500 sandbox runs to verify runtime stability, reference reproducibility, artifact generation, and scoring consistency. Tasks with unresolved scientific errors, unstable execution, evaluation misalignment, or unintended solution paths are revised or excluded, resulting in the final benchmark of 60 project-level tasks.
To measure scientific autonomy, the benchmark progressively reduces human methodological guidance. Each task is designed as a complex, project-level scientific investigation rather than an isolated question. Starting from a research objective and task-specific data, agents must carry out a long-horizon, multi-stage research process spanning problem understanding, method selection, implementation, experimentation, failure diagnosis, iterative refinement, and result validation. Across the 60 tasks, completing these research processes involves more than 2,600 interaction turns and 2,400 execution steps.
The guidance progression shifts scientific responsibility from humans to AI. In the most guided condition, the governing equations, numerical formulation, and solver procedure are explicitly provided, requiring the agent to mainly implement and execute the prescribed approach. In subsequent conditions, the full procedure is removed, leaving only methodological guidance about the class of problem and suitable approaches. In the least guided condition, the agent receives only the observed data, the scientific objective, and the required outputs, forcing it to determine the underlying model, choose an appropriate numerical method, implement it, and validate the resulting prediction. A final condition retains this minimal guidance but adds plausible yet task-irrelevant information to test whether the agent can maintain its research direction under distraction.
The benchmark covers fundamental science, life science, computing, and engineering, requiring the same agent and model system to generalize across different data types, scientific methods, and validation criteria. This breadth tests whether autonomous research transfers across disciplines rather than remaining limited to a single domain.
Experiment
ASI-Bench evaluates 18 representative Agent×Model configurations on 60 project-level research tasks across 11 scientific domains under four guidance conditions (B1–B4), where human methodological guidance is progressively withdrawn. The main results show that current systems remain far from reliable autonomous scientific discovery, with only the strongest configuration achieving a B3 score above 50, and that the sharp drop from B1 to B2 indicates the primary bottleneck is turning a selected method into a complete research procedure rather than method choice or distraction. Results also show that scientific capability emerges from the interaction between the backbone model and the agent harness. Computational-cost experiments further find that complete guidance reduces token and time costs while incomplete methodological guidance can increase overhead, and that higher spending does not reliably translate into better scientific performance.
Representative benchmarks for advanced AI capability each emphasize different aspects of autonomous research, such as broad academic knowledge, scientific coding, tool-based terminal tasks, or research replication. Cross-domain generality, method autonomy, and end-to-end research are rarely combined in a single benchmark, and none of the compared benchmarks explicitly evaluates a guidance gradient across progressively reduced methodological support. This leaves a gap for joint evaluation of general intelligence, independent method selection, and autonomous execution. Most existing benchmarks cover only one or two dimensions of autonomous research, with explicit cross-domain generality limited to a few settings. End-to-end research coverage appears in benchmarks focused on research replication, ML engineering, and AI R&D, while method autonomy is partial or absent in most others. No representative benchmark in the comparison explicitly evaluates a guidance gradient that varies the level of human methodological guidance.
Performance drops most sharply when detailed procedural guidance is removed, while removing the method choice causes a much smaller additional decline and irrelevant context has little effect. Even the strongest configuration reaches only moderate autonomous performance, though stronger inference-time reasoning offers a notable gain. Harness choice can substantially change the same model's capability, but the effect varies across model-harness pairs. The sharpest average decline comes from losing step-by-step procedural guidance, not from losing the method choice. Stronger inference-time reasoning improves autonomous method selection, but the best configuration remains moderate, and harness choices can markedly shift model scores.
The first analysis compares representative autonomous research benchmarks and finds that they rarely combine cross-domain generality, method autonomy, and end-to-end research, while none explicitly evaluates a guidance gradient with varying methodological support. The second experiment measures model performance as guidance is reduced, showing the largest drop when detailed procedural guidance is removed, a smaller additional decline when method choice is removed, and little effect from irrelevant context. Stronger inference-time reasoning improves autonomous method selection but still leaves the best configuration at moderate performance, and harness choice can substantially shift model capability.