HyperAIHyperAI

Command Palette

Search for a command to run...

L'IA dans la science : premiers enseignements

Résumé

Le progrès scientifique est un moteur clé de la croissance économique et de la prospérité. L'impact de l'IA sur la science suscite un grand enthousiasme, mais aussi des inquiétudes, bien que les données soient encore rares. Nous apportons ici des premiers enseignements à partir de trois sources de données : un échantillon de 15 millions d'interactions avec Gemini, un inventaire de plus de 2 600 modèles d'IA spécialisés couvrant diverses disciplines, et une enquête auprès de plus de 600 scientifiques. Nous cartographions ces données selon une nouvelle taxonomie des tâches scientifiques afin d'étudier comment les scientifiques utilisent l'IA. Quatre résultats principaux se dégagent. Premièrement, nous constatons une adoption et une couverture larges : les scientifiques utilisent l'IA plus que la plupart des autres professions. Les modèles d'IA spécialisés offrent une couverture disciplinaire étendue et sont fortement cités. Près de la moitié des scientifiques interrogés déclarent utiliser une forme d'IA chaque jour. Deuxièmement, nous documentons que les LLM (approximés par l'utilisation de Gemini) et les modèles spécialisés agissent comme des compléments : les LLM sont utilisés pour l'analyse générale, le codage et la préparation de manuscrits, tandis que les modèles spécialisés fournissent des prédictions spécifiques à un domaine, la génération de données et la classification. Troisièmement, les scientifiques rapportent des gains de productivité importants grâce à l'IA : une économie de près de 7 heures par semaine, temps principalement réinvesti dans davantage de recherche. Enfin, nous montrons que l'IA modifie déjà le processus scientifique. Alors que certaines étapes de la recherche deviennent plus faciles, les goulots d'étranglement se déplacent en aval. Les scientifiques signalent un arriéré croissant d'hypothèses non testées et une demande substantielle de vérification des résultats. Nos résultats suggèrent que l'IA présente un potentiel significatif pour accroître la productivité scientifique. Cependant, comme dans d'autres secteurs, son impact ultime sera déterminé par des interdépendances complexes entre les tâches et par les investissements dans l'élimination des goulots d'étranglement émergents.

One-sentence Summary

Researchers from Google, Google DeepMind, University of Chicago, CUNEF Universidad, MIT FutureTech, University of Oxford, Carnegie Mellon University, and University of Pennsylvania analyze 15×10615\times 10^615×106 Gemini interactions, over 260026002600 specialized AI models, and surveys of 600 scientists, revealing broad AI adoption across disciplines, complementarity between LLMs and specialized models, nearly 777 hours of weekly productivity savings reinvested in research, and downstream bottlenecks such as untested hypotheses and output verification demands.

Key Contributions

  • Introduces an early empirical analysis of AI adoption in science using three complementary data sources: 1.5 million Gemini interactions, an inventory of over 2,600 specialized AI models, and a survey of more than 600 scientists, all mapped to a new taxonomy of scientific tasks.
  • Documents a complementary division of labor between general-purpose LLMs and specialized models, showing that LLMs handle general analysis, coding, and manuscript preparation while specialized models provide domain-specific predictions, data generation, and classification.
  • Quantifies productivity gains and emerging bottlenecks: scientists report saving nearly 7 hours per week, reinvesting most of that time into research, yet nearly 90% of those users spend a meaningful share of saved time verifying AI outputs, and a growing backlog of untested hypotheses indicates downstream workflow constraints.

Introduction

Scientific progress has long been a driver of economic growth, but recent concerns about declining research productivity have sparked interest in whether AI can reverse this trend. While specialized models like AlphaFold have shown Nobel-caliber potential, the debate over AI's impact on science has relied mostly on lagging indicators like publications and patents, with little real-time evidence on how researchers actually use these tools.

The authors address this gap by combining three complementary data sources: 15 million Gemini interactions, bibliometrics from over 2,600 specialized AI models, and a survey of more than 600 scientists. They map all data to a new taxonomy of scientific tasks developed by MIT FutureTech, enabling task-level analysis across academic disciplines. Prior work lacked this granularity, as existing taxonomies like O*NET flatten the domain-specific activities that define research, and frontier AI lab telemetry had not been focused on scientific workflows.

The main contribution is an empirical, multi-source snapshot of AI adoption in science. The analysis reveals that scientists lead other occupations in AI usage, that LLMs and specialized models act as complements with a clear division of labor, and that while researchers report substantial time savings, bottlenecks persist in physical experimentation and validation. The authors also document a concerning trend: nearly half of surveyed scientists say AI pushes them toward safer, incremental projects rather than high-risk discoveries, highlighting a potential "streetlight effect" where AI lowers the cost of incremental work without advancing the scientific frontier.

Dataset

Dataset Composition and Sources

The authors combine five complementary data sources to evaluate AI in science:

  • Google ATLAS 1.0 corpus: approximately 15 million anonymized user interactions from the Gemini App, Google AI Mode, and API interfaces, sampled in early April 2026.
  • Specialized AI for Science Model Inventory: a curated set of 2,690 notable AI models for science, assembled through agentic search across scientific repositories and literature, augmented by Epoch AI's Foundation Model database, and enriched with publication, citation, and institutional metadata from OpenAlex.
  • Scientist Survey: a survey of 637 active scientists in the US and UK, run by a third-party provider through the More in Common online panel between 27 July and 11 August 2026.
  • MIT FutureTech Scientific Task Taxonomy: built from approximately 3.8 million Lightcast (2026) scientific research job adverts posted after 2010.
  • OpenAlex Publication Database: bibliometric records providing metadata on scientific works, including publication dates, author affiliations, funding sources, citations, and concept classifications.

Key Details for Each Subset

  • Gemini Interaction Logs: The authors isolate scientific use cases through a three-stage filtering pipeline. First, they discard non-work interactions using ATLAS's work versus non-work classifier. Second, they apply ATLAS's occupation predictions, restricting the sample to six 2-digit Standard Occupational Classification (SOC) categories covering physical, life, social, computer, engineering, and health disciplines. Third, they apply a custom "science" classifier that operationalizes their definition of science. This yields about 360,000 science interactions from the initial 15 million.
  • Specialized Model Inventory: The inventory includes domain-specific AI models (protein structure predictions, genomic models, neural climate models) and deep learning models (neural networks, diffusion models, vision transformers) that could help scientific discovery. The authors exclude general statistical ML (regularized regression, random forests) and general software or ML frameworks (PyTorch, TensorFlow). Models were published after 2012 and linked to a code repository. The inventory covers all 26 scientific fields in the OpenAlex taxonomy and 83% of the subfields.
  • Scientist Survey: Participants were identified as scientists according to four criteria, including that their main job involves working directly in science and technology, clinical or health research, life sciences, or social science research. The survey collects information on demographics, fields of research, AI adoption intensity, time allocation across research tasks, expected time savings from AI, and perceived impacts on interdisciplinarity and bottlenecks.
  • MIT FutureTech Scientific Task Taxonomy: The pipeline extracts over 64 million task instances from scientific job adverts, consolidates semantically equivalent instances into approximately 210,000 representative tasks, and organizes them under a global taxonomy of 12 Level-1 areas, 114 Level-2 areas, and 2,433 Level-3 groups. The hierarchy is built using hierarchical clustering of high-dimensional task embeddings and LLM analysis.
  • OpenAlex: The authors audit and clean metadata to address known issues in OpenAlex classifications, such as its tendency to disproportionately assign computational and data-driven papers under Computer Science.

How the Paper Uses the Data

  • Classification and Mapping: The authors use the OCTO pipeline, which integrates vector embeddings, clustering algorithms, and frontier Gemini LLMs, to map unstructured text into structured representations. OCTO classifies both Gemini interaction logs and publication abstracts against the MIT FutureTech Scientific Task Taxonomy and OpenAlex disciplinary hierarchy.
  • Taxonomy Filtering: The authors focus analysis on 217 scientific subfields that have research activities across both commercial and non-commercial sectors, covering 92% of all OpenAlex works and 96.5% of citations. They exclude non-R&D subfields and aggregate others where domain boundaries were not sharp.
  • Task Extraction: For specialized models, the authors use OCTO to extract 4,475 tasks from publication abstracts, formatted using the same verb / object / context structure as the MIT FutureTech Scientific Task Taxonomy. They map these tasks to the taxonomy to understand what tasks and fields could benefit from specialized model outputs.
  • Survey Analysis: The survey data is used to assess broader economic and organizational impacts of AI, including bottlenecks in the research process, verification costs, and perceived impacts on research question ambition and quality.

Processing Details

  • Science Definition: The authors adopt a broad definition of science, including academic, industry, non-profit, and government roles, and define science as any pursuit likely to generate new knowledge through investigation, computation, experimentation, or theoretical work.
  • Occupational Over-Representation: The authors calculate an occupational over-representation ratio, comparing an occupational group's share of Gemini conversation volume to its baseline share of US employment. Research-intensive occupational families show about 1.8 times higher Gemini interaction volume than their labor force share.
  • Science Interaction Characteristics: Compared to average work conversations, science interactions are about 7% more multimodal, 19% higher in token usage, 11% higher in number of turns, and measure 26% higher domain expertise score.
  • Task Coverage: The authors find that 17.6% of Level-3 tasks in the taxonomy are covered by models in the inventory, rising to 24.3% when excluding operational and teaching tasks unlikely to be mentioned in publication abstracts.

Method

The authors leverage a comprehensive methodology to quantify scientific engagement with AI, utilizing two distinct data streams and a unified classification framework.

Gemini Interaction Logs Filtering Pipeline To isolate scientific use cases from the Google ATLAS 1.0 corpus, which contains approximately 15 million anonymized user interactions, the researchers implement a three-stage filtering pipeline. The first stage employs a work versus non-work classifier to discard interactions related to personal, leisure, or educational inquiries. In the second stage, the sample frame is restricted to six Standard Occupational Classification categories where research activities concentrate, including physical, life, social, computer, engineering, and health disciplines. The final stage applies a custom science classifier to the content and context of the interactions, identifying those likely to represent scientific workflows. This process yields a corpus of approximately 360,000 science-specific interactions.

The authors then map this corpus using the OCTO classification framework across the OpenAlex disciplinary hierarchy and the MIT FutureTech Scientific Task Taxonomy. As shown in the figure below, the distribution of these interactions across high-level task categories indicates a heavy concentration in quantitative data analysis and modeling.

Specialized AI Model Inventory Construction Alongside general-purpose LLMs, the authors construct a curated inventory of 2,690 specialized AI models for science. This inventory is assembled through agentic search across scientific repositories and literature, augmented by the Epoch AI Foundation Model database, and enriched with publication and citation metadata from OpenAlex. The selection includes domain-specific models and deep learning architectures while excluding general statistical machine learning and software frameworks.

The distribution of these models across scientific domains is broad. As illustrated in the framework diagram, Physical Sciences and Computer Science dominate the inventory, though Life and Health Sciences also account for a significant portion of the models, particularly in fields like Biochemistry, Genetics, and Molecular Biology.

Task Extraction and Taxonomy Mapping To evaluate the downstream capabilities of the specialized models, the authors utilize the OCTO semantic classification tool to extract tasks from the abstracts of canonical publications. These tasks are formatted using a verb, object, and context structure and mapped to the MIT FutureTech Scientific Task Taxonomy. The extraction process identifies over 4,475 tasks associated with more than 200 unique actions.

The analysis of action verbs reveals distinct usage patterns across different scientific fields. Specialized models are primarily utilized for prediction, generation, and classification, though field heterogeneity reflects specific use cases, such as segmentation in Medicine or solving mathematical proofs in Mathematics.

When mapping these extracted tasks to the MIT FutureTech Scientific Task Taxonomy, the authors find that nearly one in four substantive research tasks are covered by the capabilities in the inventory. The most common task category at the highest level of the taxonomy is the analysis and modeling of quantitative data, which aligns with the primary role of specialized models in predicting and simulating data for downstream analysis.

Experiment

The experiments analyze AI adoption in science by combining Gemini interaction logs, a specialized model inventory, and survey data. Results show LLM usage spans nearly all scientific subfields, led by physical and computer sciences, with usage scaling proportionally to researcher populations. Task analyses reveal that while both general LLMs and specialized models concentrate on data analysis, they operate as complements at granular levels, with specialized models handling domain-specific prediction and simulation. Survey data corroborate this picture, showing that most researchers use AI weekly and report net time savings reinvested into higher output, though verification of AI outputs and task bottlenecks temper productivity gains.

LLM adoption in science varies notably across disciplines, with each field showing distinct over- and under-represented task types relative to average researcher activity. Health and Life Sciences lean toward clinical and experimental tasks, while Computer Science and Physical Sciences favor software, data analysis, and prototyping tasks. Health Sciences shows nearly five times higher relative use for clinical study coordination, with compliance and regulatory documentation also overrepresented. Life Sciences has over three times higher relative use for experimental and sample processing tasks. Computer Science is overrepresented in software troubleshooting and quantitative data analysis, while Physical Sciences is overrepresented in prototype and process technologies. Social Sciences is overrepresented in data collection, teaching, and project management, but underrepresented in experimental and product development tasks. Clinical study coordination is consistently underrepresented in Computer Science, Physical Sciences, and Life Sciences, while experimental operations are underrepresented in Computer Science and Social Sciences.

An analysis of LLM adoption across scientific disciplines reveals distinct task-level specialization patterns. Health and Life Sciences show strong overrepresentation in clinical, experimental, and compliance-related tasks, while Computer Science and Physical Sciences favor software, data analysis, and prototyping. Social Sciences are overrepresented in data collection, teaching, and project management but underrepresented in experimental work. Notably, clinical study coordination is consistently underrepresented in Computer Science, Physical Sciences, and Life Sciences, whereas experimental operations are underrepresented in Computer Science and Social Sciences.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp