Google Unveils Gemini 4 Argon, Retaking Benchmark Lead Over OpenAI and Anthropic (venturebeat.com) 1
Google has unveiled Gemini 4 Argon, a new frontier AI model that it says leads or ties rivals on 13 of 18 disclosed benchmarks. The model is initially being released only to trusted cyber defenders and select pre-release testers, with broader availability planned later. VentureBeat reports: Google is not claiming that Gemini 4 Argon wins every benchmark. But across the benchmark table disclosed in the company's embargoed materials, Argon posts the highest score, or ties for the highest score, in more categories than GPT-6 Astra or Claude Opus 5.5. Across the 18 benchmarks Google disclosed, Argon leads outright on 12 and ties for first on one. GPT-6 Astra leads outright on three and ties Argon on one. Claude Opus 5.5 leads outright on two. That makes Argon the leading frontier model by total number of benchmark leads or top scores in Google's comparison set, even though the results show a still-close race in which OpenAI and Anthropic retain advantages in several important technical categories.
[...] The result is a more nuanced claim than "best model across the board." Argon appears to have the broadest top-score profile among the three frontier models in Google's disclosed comparison, but GPT-6 Astra remains ahead in several software, science-terminal and computer-use tasks, while Claude Opus 5.5 remains ahead in terminal-agent and post-training workflows. For enterprise buyers, that means model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count. [...] Argon also expands Google's output ceiling. The company says the model supports an industry-leading 1 million output tokens, up from a previous 64,000-token limit. That is a notable distinction for agentic software engineering, audit, migration and legal-review workloads where the value of a model often depends on how long it can sustain a chain of work before handing control back to a human.
[...] The result is a more nuanced claim than "best model across the board." Argon appears to have the broadest top-score profile among the three frontier models in Google's disclosed comparison, but GPT-6 Astra remains ahead in several software, science-terminal and computer-use tasks, while Claude Opus 5.5 remains ahead in terminal-agent and post-training workflows. For enterprise buyers, that means model choice is still workload-dependent, even if Argon now gives Google its strongest claim yet to overall frontier leadership by benchmark count. [...] Argon also expands Google's output ceiling. The company says the model supports an industry-leading 1 million output tokens, up from a previous 64,000-token limit. That is a notable distinction for agentic software engineering, audit, migration and legal-review workloads where the value of a model often depends on how long it can sustain a chain of work before handing control back to a human.