OpenAI Launches GPT-6 Astra, Achieves Near-Perfect Benchmark Scores
OpenAI officially released GPT-6 Astra on September 3, positioning the model as its most capable and aligned system to date. During the launch, President Greg Brockman declared, "Welcome to the AGI era." The model is initially available to a select group of institutions, with broader access scheduled for ChatGPT Plus, Pro, Business, Enterprise, and API users in the coming days. Enterprise administrators must manually enable the model, which will be deactivated by default upon initial release. GPT-6 Astra introduces a significant leap in autonomous task execution. Capable of managing complex, multi-step workflows spanning browser navigation, document analysis, code generation, and professional software operations, the model autonomously sequences steps, adapts to errors, and maintains persistent work notes across sessions. This architecture dramatically improves efficiency, cutting average task completion time by approximately 47 percent compared to its predecessor. Performance benchmarks reflect these gains: Astra scored 72.6 percent on OSWorld 2.0, 41.4 percent on AutomationBench, and 57.7 percent on Terminal-Bench 4.0. In advanced reasoning tests, it achieved 97.6 percent on FrontierMath Tier 4 and 99.9 percent on the ARC-AGI-3 benchmark, surpassing the previous generation by over ninety percentage points. Technically, the model features a one-million-token context window, a maximum single output of 128,000 tokens, and supports async tool calling, allowing concurrent processing while awaiting tool responses. Early adopters report substantial workflow improvements. Legal technology firm Legora noted a near-40 percent performance increase in financial document auditing, while game developer Playco reduced manual debugging by half and accelerated prototype generation. OpenAI has adjusted the pricing structure accordingly, setting API rates at 10 dollars per million input tokens and 50 dollars per million output tokens. While these figures represent a 2.5 times increase over prior models, the company emphasizes that reduced step counts and output lengths lower the effective cost per completed task. The advancement also necessitates robust safety measures. Astra is the first OpenAI model to achieve the "Critical" cybersecurity rating, demonstrating 100 percent efficacy on ExploitBench and identifying two previously unknown zero-day vulnerabilities. In response, standard deployments will block high-level exploit code generation, and all tool-use reasoning processes will undergo enhanced monitoring, addressing the increased opacity of autonomous self-reflection.
