Gemini 3.1 Pro Outperforms Competitors in Reasoning and Safety Benchmarks with Strong Ethical Safeguards
Gemini 3.1 Pro demonstrates strong performance across a wide range of benchmarks, particularly in reasoning, coding, and multimodal understanding. On Humanity's Last Exam, a comprehensive academic reasoning test, Gemini 3.1 Pro achieves 44.4% without tools and 51.4% with search and code capabilities, outperforming Gemini 3 Pro and Sonnet 4.6, though falling short of GPT-5.2 and GPT-5.3-Codex in some areas. In abstract reasoning, it scores 77.1% on ARC-AGI-2, significantly ahead of Gemini 3.0 Pro and Sonnet 4.6, though below GPT-5.2 and GPT-5.3-Codex. For scientific knowledge, Gemini 3.1 Pro excels with 94.3% on GPQA Diamond, surpassing all other models listed and showing strong consistency in factual accuracy. In coding tasks, it leads on SWE-Bench Verified with 80.6% on single attempts, closely matching Sonnet 4.6 and outperforming GPT-5.2. It also performs well on LiveCodeBench Pro and SciCode, achieving 59% and 59% respectively, indicating solid capabilities in scientific and competitive programming. On long-horizon professional tasks, Gemini 3.1 Pro scores 33.5% on APEX-Agents, outperforming most models except Sonnet 4.6. In expert-level task evaluation, it achieves an Elo rating of 1317, ranking above GPT-5.2 and GPT-5.3-Codex. For multi-step workflows using MCP, it scores 69.2%, demonstrating strong agentic planning and execution. In agentic search tasks, it achieves 85.9% on BrowseComp, showing strong integration of browsing and code execution. In multimodal reasoning, Gemini 3.1 Pro scores 80.5% on MMMU-Pro and 92.6% on MMMLU, outperforming Sonnet 4.6 and GPT-5.2 in key areas. It maintains strong long-context performance, scoring 84.9% on MRCR v2 at 128k tokens, and remains competitive at 1M tokens despite a drop to 26.3%. In safety evaluations, Gemini 3.1 Pro shows a slight improvement in text-to-text and multilingual safety, with a +0.10% and +0.11% increase in safety policy adherence. However, image-to-text safety declined slightly by -0.33%. The model maintains a neutral tone and reduced unjustified refusals. The Frontier Safety Framework assessment confirms that Gemini 3.1 Pro remains below alert thresholds for critical capability levels (CCLs) across CBRN, harmful manipulation, machine learning R&D, and misalignment. While cyber capabilities have increased compared to Gemini 3.0 Pro, the model still does not reach the CCL threshold. Deep Think mode does not significantly boost performance in high-cost inference scenarios. The model shows improved capabilities in machine learning R&D, particularly on RE-Bench and Optimise LLM Foundry, but remains below the alert threshold. On misalignment and instrumental reasoning, performance is consistent or slightly improved, but no CCLs are reached. Overall, the model maintains a strong safety profile with ongoing mitigations in place.
