HyperAIHyperAI

Command Palette

Search for a command to run...

AI Defeats Top Stratego Players With Efficient Training and On-the-Fly Planning

Researchers from MIT, Carnegie Mellon University, New York University, and Stanford University have developed Ataraxos, an artificial intelligence system that achieves superhuman performance in Stratego, a complex war game defined by hidden information. The breakthrough, detailed in a recent study published in Nature, demonstrates a highly efficient approach to decision-making under uncertainty, with potential applications in military strategy, cybersecurity, and high-stakes negotiations. Stratego presents a significant challenge for AI due to its vast number of possible piece configurations, exceeding 10 to the 66th power. Unlike perfect-information games such as chess, where optimal moves remain constant regardless of context, Stratego requires constant estimation of an opponent’s concealed assets and strategic bluffs. Previous AI attempts, including DeepMind’s DeepNash, required massive computational resources and still fell short of defeating elite human competitors. To overcome these limitations, the research team engineered Ataraxos using a dual-method architecture. The system initially trains through self-play reinforcement learning, rapidly developing a foundational blueprint strategy while avoiding exhaustive calculation of every possible scenario. This approach drastically reduces computational overhead, requiring less than one percent of the training data and one thirtieth of the self-play iterations needed by prior models. During active gameplay, Ataraxos refines its decisions in real time using decision-time planning. A specialized generative model continuously estimates the probability distribution of the opponent’s hidden pieces, allowing the system to evaluate future board states and select optimal moves tailored to the specific match context. The resulting system outperformed all existing AI benchmarks and decisively beat top human competitors. In head-to-head matches, Ataraxos recorded a 15–1–4 victory against the world’s leading Stratego player and secured a 39–2 record against elite competitors at the annual Stratego World Championship. The AI demonstrated superior risk assessment, maintaining composure and avoiding the overcorrection often seen in human players when high-value units are threatened. Beyond its success in standard Stratego, the framework proved highly adaptable. Researchers successfully deployed the same core methodology to Barrage Stratego, a faster variant, as well as cooperative and competitive multi-player games like Hanabi and Dou dizhu, achieving superhuman results across all tested environments. This versatility underscores the system’s capacity to generalize across different imperfect-information domains. Senior author Gabriele Farina of MIT emphasized that real-world strategic environments rarely offer complete data transparency, making scalable AI decision-makers increasingly vital. While Ataraxos currently operates as a highly effective black box, the team plans to integrate interpretability measures to allow human auditors to understand and verify the AI’s reasoning before broader operational deployment. The research was led by Samuel Sokota of Carnegie Mellon University, with contributions from researchers at NYU, Stanford, and MIT, marking a significant step forward in algorithmic strategy under uncertainty.

Related Links