Google on Wednesday announced Gemini 4 Argon, the company’s long-awaited flagship model, and it looks like it was worth the wait.Across most benchmarks, Gemini 4 Argon beats OpenAI’s and Anthropic’s top models, sometimes by a wide margin, though where it trails, it can trail by as much as 10 points.Google notes it is taking a phased approach and has “actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access.”Gemini 4 Argon beats OpenAI’s and Anthropic’s top models by a wide margin.The announcement comes only a day after Google CEO Sundar Pichai co-signed a commitment to “self-police” after a meeting with President Donald Trump. Anthropic, Meta, Nvidia, OpenAI, and SpaceX also signed this “commitment,” which doesn’t seem to come with any enforcement mechanism.CategoryBenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5Knowledge workVals Index68.9%63.1%65.8%67.0%Knowledge workAutomationBench (Score)51.3%41.4%31.4%42.5%Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%Knowledge workHarvey’s Legal Agent Benchmark19.6%5.4%6.7%3.8%Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%Agentic codingTerminal-bench 4.057.4%58.2%57.9%66.4%ML engineeringPostTrainBench45.3%44.3%40.2%49.3%Science and mathTerminal-Bench Science 0.157.6%68.1%52.6%63.3%Science and mathRiemannBench76.0%72.0%65.6%69.6%Long contextGraphWalks (Up to 128k, BFS (F1))99.7%98.7%91.4%90.6%Long contextGraphWalks (256k to 1M, BFS (F1))84.2%71.8%65.0%66.8%Computer useAgent’s Last Exam (Pass rate)39.5%34.2%—38.2%Computer useOSWorld-2.0 (Offline subset, Partial reward)69.2%72.6%——Multimodal understandingChartography71.6%71.0%46.2%66.3%Multimodal understandingLVBench91.7%87.5%79.7%83.7%CybersecurityCWE-bench v168.0%68.0%58.0%67.0%Google says it will gather feedback from early testers and iterate on Argon’s guardrails before making the model available outside the Fairwind Program. Once that day comes, paid API customers and AI Ultra subscribers will get to try the new model first, before it’s released to developers, enterprises, and consumers.Unlike its boring, odorless, and inert namesake, Gemini 4 Argon will likely create a bit of a stir. Google first announced plans for a new Pro model, Gemini 3.5 Pro, at its I/O developer conference in May — Argon is essentially its replacement. The original plan was to launch the new model in June, but instead, we got a series of Flash models.Coding: so-so. Knowledge work: A+In Google’s benchmarks, which include a wide variety of tasks, Argon takes top billing, outright or tied, in 13 of 18 tests against Anthropic’s Opus 5.5 and Fable 5.1, as well as OpenAI’s GPT-6 Astra.Yet while Google specifically mentions Argon’s coding abilities, the results are mixed.Google highlights Argon’s 77.9% on DeepSWE v1.1 as a new state of the art, but Argon also comes last among the competition on both FrontierSWE v2 and Terminal-Bench 4.0, where GPT-6 Astra and Opus 5.5 lead it by 10.5 and nine points, respectively. Its other coding win is on Vibe Code Bench, with 91.9%, but all four models score above 89% here.But where Argon excels is knowledge work.Argon scores 51.3% on Zapier’s AutomationBench, almost nine points ahead of Opus 5.5, and 84.2% on the GraphWalks test for inputs between 256K and 1M tokens, more than 12 points ahead of GPT-6 Astra. Its 19.6% on Harvey’s Legal Agent Benchmark is nearly triple Fable 5.1’s score, though that means it still fully completes only about one in five tasks.In other areas, Argon’s wins are narrower. While it leads on the Vals Index, Vibe Code Bench, Agent’s Last Exam, Chartography, and the shorter GraphWalks test, those leads are generally under two points.On CWE-bench v1, the one cyber benchmark in the post with rival scores, Argon ties GPT-6 Astra and xAI’s Grok 4.7 at 68%, with Opus 5.5 a point behind, but it’s worth noting that the OpenAI and Anthropic models run in their own agent harnesses (Codex and Claude Code), so that leaderboard measures each model and its tooling together.As always, benchmarks never tell the full story, but it looks like Google focused on making this model especially useful for standard office tasks.Developers will surely want to test the model, too, given the impressive DeepSWE score. 1 million output tokensOne interesting new feature in an area where most of the competition isn’t currently pushing the frontier forward is in the model’s output token limits. One million input tokens is now the standard for frontier models, but Gemini 4 Argon can also generate up to one million output tokens, up from 64,000 for previous Gemini models.“When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go,” Google writes in the announcement.As for cyber security, Google says it trained Argon to autonomously find, validate, and patch software vulnerabilities, and for Fairwind participants and its own internal teams, it’s releasing the model without cyber guardrails. Wiz, which Google acquired for $32 billion in March, is already using Argon in its Scan for Good initiative, and Google says the model found a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models had missed. Argon scores 85.8% on Google’s internal vulnerability discovery benchmark and 70.9% on Wiz’s penetration testing benchmark, though Google only compares those results with its own Gemini 3.8 Flash Cyber (71.0% and 58.2%, respectively). Pricing?Google says Argon will cost $2 per million input tokens and $10 per million output tokens during an introductory period, rising to $4 and $20 afterward. That later output rate matches the $20 per million output tokens Anthropic charges for Opus 5.5, and with ArgonThe post Gemini 4 Argon is here: It’s great, and you can’t have it yet appeared first on The New Stack.