Command Palette
Search for a command to run...
LLM은 도약할 수 없다
LLM은 도약할 수 없다
Tom Zahavy Oliver Nash
초록
우리는 근본적으로 새로운 것을 어떻게 발견하는가? 알베르트 아인슈타인은 모리스 솔로빈에게 보낸 편지에서 발견을 감각 경험에서 공리로의 직관적 ‘도약’과 그에 뒤따르는 논리적 연역으로 구성된 순환적 과정으로 개념화했다. 생성형 AI가 귀납(통계적 패턴 매칭)을 숙달하고 연역(형식적 증명)을 빠르게 정복해 가고 있지만, 우리는 AI에 새로운 설명적 가설을 생성하는 귀추법의 메커니즘이 결여되어 있다고 주장한다. 아인슈타인의 일반 상대성 이론 정립 과정을 계산적 사례 연구로 삼아, ‘데이터 압축으로서의 창의성’(귀납)이라는 지배적인 이론이 관측 데이터가 희소한 발견을 설명하지 못함을 입증한다. 본 입장 논문은 현대의 대규모 언어 모델이 확립된 전제로부터 정리를 증명하는 연역 단계는 그럴듯하게 수행할 수 있겠지만, 그러한 전제를 정식화하는 데 필요한 귀추적 ‘도약’을 구조적으로 수행할 수 없음을 논증한다. 우리는 시뮬레이션을 형식적 공리로 변환하는 과정을 인공적 과학적 발명의 결정적 병목 지점으로 식별하며, 물리적으로 일관된 다중 모달 세계 모델이 이 간극을 메우는 데 필요한 감각적 토대를 제공한다고 제안한다.
One-sentence Summary
Google DeepMind researchers argue that large language models, though adept at induction and deduction, are structurally incapable of the abductive “jump” necessary to formulate novel explanatory hypotheses, and propose that physically consistent multimodal world models can provide the sensory grounding required to overcome this critical bottleneck in artificial scientific invention.
Key Contributions
- The paper challenges the "creativity as compression" hypothesis by demonstrating that Einstein's formulation of General Relativity occurred without a pervasive error signal or abundant observational data, showing that statistical induction alone cannot explain such discoveries.
- This work identifies the abductive leap, translating internal physical simulations into formal axioms, as the critical missing mechanism in current AI, and supports this claim by examining the sensory-grounding limitations of state-of-the-art automated discovery systems like the AI Scientist and AlphaEvolve.
- It proposes that physically consistent, multimodal world models can provide the necessary sensory grounding for AI to perform the counterfactual simulations that drive abduction, offering a pathway to bridge the gap between intuition and logic in physical-science invention.
Introduction
The authors challenge the reductionist view that scientific discovery is merely the combination of induction (data compression) and deduction (formal theorem proving), using Einstein's creation of General Relativity as a computational case study. Prior work in AI-driven discovery has shown success in extracting patterns from data or verifying theories from given axioms, but it cannot explain how novel axioms are formulated when no error signal or large dataset exists. The authors argue that the critical missing ingredient is abduction—the creative inference of a cause to explain a surprising phenomenon—and demonstrate that Einstein's axiomatic leap was driven by embodied thought experiments, not by statistical gradient. They conclude that current large language models, which lack physically consistent world models, are structurally incapable of the abductive jump required for true scientific invention, and propose that future AI must integrate sensory simulation to ground abstract symbols in physical experience.
Method
The authors propose a conceptual framework for automating scientific discovery that bridges the gap between logical deduction and the generation of novel axioms. They argue that while modern AI systems excel at deduction, which involves deriving theorems from existing axioms, they lack the capacity for the abductive jump required to formulate new foundational principles. To address this, the authors introduce the concept of Manipulative Abduction, a cognitive process relying on embodied simulation to generate hypotheses through active interaction with mental models.
This thought process is conceptualized as a two-stage mechanism. The first stage involves Simulation as Physical Variation, which can be viewed through the lens of Test Time Reinforcement Learning. In this phase, the system invents new variations of a problem to force progress. Unlike symbolic variations bounded by fixed rules, this requires manipulating perceptual experience to invent new axioms based on physical intuition.
As illustrated in the figure above, the equivalence principle thought experiment exemplifies this stage. By simulating the physical sensations of an observer inside a sealed environment, such as comparing a stationary elevator in a gravitational field to an accelerating elevator in deep space, the system generates a specific sensory pattern. The second stage is Abduction and the Physical Prior, where the system performs abductive reasoning to infer the best explanation for the simulated observation. Because the simulated sensory experience of acceleration is indistinguishable from the remembered sensory experience of gravity, the system abducts that they must be the same phenomenon, effectively pruning the search space of possible axioms using a physical prior.
To operationalize this in AI, the authors point to the emergence of interactive World Models. Unlike passive video generators that rely on statistical correlation, architectures like Genie introduce action controllability, allowing for agentic intervention. This capability is a prerequisite for Manipulative Abduction, as it enables the AI to perform counterfactual interventions, such as conceptually cutting the cable of an elevator, rather than merely observing. By operating on a consistent latent physics manifold, these interactive environments provide a synthetic laboratory to transform the abductive jump from a mystical insight into a reproducible algorithmic process.