DeepMind Safety Researcher Resigns for Independent AI Evaluation
A growing cohort of researchers from leading artificial intelligence laboratories is departing for independent evaluation organizations, driven by urgent concerns over the safety of recursive self-improvement and the potential for catastrophic alignment failures. On September 13, Josh Engels, a researcher from Google DeepMind's AGI safety team, announced his transition to the non-profit Model Evaluation and Threat Research (METR), effectively declining recruitment offers from OpenAI and Anthropic. This move follows a similar announcement on September 11 by Joe Benton, former head of Anthropic's scalable oversight team, who also joined METR. Engels, whose research background encompasses mechanical interpretability and scalable oversight, warns that the industry's pursuit of superintelligence involves AI systems increasingly designing their successors. He describes a dangerous feedback loop where models participate in their own development, risking the amplification of alignment errors as capabilities accelerate. Engels predicts a "dreaded possibility" of significant harm within five years, emphasizing that current methods may be insufficient to verify safety once models surpass human comprehension. Benton shares these apprehensions, noting that companies are racing toward an intelligence explosion without allocating commensurate resources to safety, leaving the public unaware of critical progress and near-misses. METR, formerly ARC Evals, serves as an independent body dedicated to testing frontier AI capabilities and risks without regulatory authority but with the mandate to validate safety claims externally. The agency focuses on measuring autonomous task completion, detecting evasion behaviors, and investigating accidents to build public evidence of industry risks. Engels plans to investigate the origins of misalignment behaviors during training and assess whether current mitigation strategies are adequate, while Benton aims to advocate for mandatory disclosure of recursive self-improvement progress and safety incidents by independent labs. The exodus highlights a structural tension in AI development: competitive pressure drives rapid advancement, making voluntary safety pauses difficult. While independent evaluation provides essential external scrutiny, it faces limitations, including reliance on company-provided data and the lack of enforcement power. Critics note that defining safety metrics can also serve to entrench incumbent advantages, yet the consensus among these researchers is that external verification is critical to ensuring that alignment progress keeps pace with model capability.
