HyperAIHyperAI

Command Palette

Search for a command to run...

Xiaomi Scales Robot Foundation Models With 100,000 Hours of Real Data

Xiaomi has advanced the frontier of robotic intelligence with the release of Xiaomi-Robotics-1, a novel vision-language-action model engineered to overcome historical data scarcity in robot policy training. The research addresses a fundamental bottleneck in the field: the lack of large-scale, high-quality training datasets that have traditionally capped the scalability of robotic foundation models. To bridge this gap, the company introduced a two-stage training paradigm inspired by large language model development, combining massive embodiment-free pre-training with targeted post-training alignment. The model architecture rests on 100,000 hours of embodiment-free trajectory data, automatically annotated across more than 1,700 diverse scenarios spanning household, commercial, industrial, and outdoor environments. This initial pre-training phase enables the system to learn generalized action representations and scene transition dynamics without being constrained by specific hardware configurations. During the subsequent post-training stage, the model is fine-tuned using 7,200 hours of real-world robot data collected from actual domestic settings, alongside curated open-source and manually annotated datasets. This hybrid approach aligns the policy network with physical robot constraints while preserving its capacity to follow complex natural language instructions. Xiaomi-Robotics-1 demonstrates remarkable data efficiency and generalization capabilities across downstream applications. The system can master highly complex, novel real-world tasks using fewer than ten hours of demonstration data per task, achieving a 75 percent overall success rate. This performance nearly doubles that of competing baselines under equivalent data budgets. When training data is expanded to under forty hours per task, the success rate climbs to 85 percent. Beyond real-world adaptation, the model establishes new state-of-the-art benchmarks across four major simulation suites, including RoboCasa, VLA-Bench, and RoboDojo, with relative performance gains reaching over fifty percent against secondary competitors. By decoupling initial representation learning from physical embodiment constraints, Xiaomi’s methodology offers a scalable pathway for training robust robot policy models. The findings, detailed in a recent technical preprint, underscore the viability of leveraging massive cross-embodiment datasets to accelerate the development of general-purpose robotic systems. This approach signals a pivotal shift in how AI-driven automation can be trained, optimized, and deployed across varied operational environments, moving the robotics industry closer to widely adaptable, data-efficient machine intelligence.

Related Links

Unknown SourceUnknown Source