Reflection Unveils Beam 501B Open-Weight Model Rivaling Chinese Models
Reflection AI, a Brooklyn-based startup founded in 2024 by former Google DeepMind researchers, has officially unveiled Beam, its inaugural open-weight artificial intelligence model. Designed as a competitive Western alternative to leading Chinese open-source models and frontier closed systems, Beam operates as a 501-billion-parameter sparse Mixture-of-Experts architecture with 23 billion active parameters and a one-million-token context window. The company plans to release the full weights, technical report, and developer tools under an Apache 2.0 license later this month, following an early access phase for select users. Beam capabilities stem from an extensive pretraining phase and a high-compute reinforcement learning pipeline. The base model was trained on 23.8 trillion curated tokens sourced from web content, proprietary datasets, and technical repositories. To enhance reasoning and agentic performance, Reflection deployed 10,500 NVIDIA GB300 GPUs for a four-week training cycle that generated over 100 million rollouts across nearly one million synthetic and proprietary environments. The company reports that Beam matches the advanced reasoning and coding benchmarks of Z.AI GLM-5.2 while consuming three to four times less inference compute. Independent verification of these performance claims remains pending. The startup has positioned Beam as a cost-efficient workhorse for enterprise, public sector, and developer workflows. This strategic focus aligns with Reflection broader ambitions to establish sovereign AI infrastructure. Backed by approximately 4.7 billion dollars in funding and recently securing over 7 billion dollars in compute agreements with SpaceX and Nebius for NVIDIA next-generation chips through 2029, the company is pursuing a vision of customizable AI factories. These systems would enable institutions to fine-tune Reflection foundation models on proprietary data, catering to sectors ranging from hedge funds to international government partnerships, including early trials with South Korea Shinsegae Group. Beam technical architecture emphasizes stable optimization and efficient token usage. Reflection implemented asynchronous policy gradients, novel load-balancing mechanisms for expert routing, and a controllable reasoning-effort parameter that allows users to balance response length against computational overhead. Safety and alignment were integrated via a multi-teacher on-policy distillation framework, combining reinforcement learning with deliberate alignment techniques to mitigate hallucinations and ensure policy compliance. With its immediate release across hyperscalers and open-source libraries, Reflection aims to strengthen the Western open-weight ecosystem. Beam marks the first iteration in the company model series, with subsequent versions already in development to narrow the gap between open and frontier intelligence.
