Google just put a number on the table that OpenAI and Anthropic are going to have to answer for. The new model is called Gemini 4 Argon, and depending on which of eighteen disclosed benchmarks you look at, it either leads or ties the field on thirteen of them. That is the kind of stat Google’s marketing team dreams about. It is also, if you squint at the fine print, a little less dominant than the headline suggests.
What Argon Actually Does Better
Argon’s strongest showings aren’t in the areas most people associate with “smartest AI model.” It crushes the competition on Harvey’s legal-analysis benchmark, posting 19.6 percent accuracy against rivals stuck in the 3.8 to 5.4 percent range. That is not a typo, and it is not a close race. The model also leads on AutomationBench, a measure of how well an AI can execute multi-step business workflows without a human nudging it along, and on GraphWalks, which tests long-context reasoning across sprawling documents.
Coding is where things get interesting. Argon posts a strong 77.9 percent on DeepSWE v1.1, ahead of where Google’s previous models landed, and it also leads on financial analysis (65.4 percent on the Vals Finance benchmark) and video understanding (91.7 percent on LVBench). For a model aimed squarely at enterprise customers, that spread is deliberate. Google isn’t trying to win the “write me a poem” contest. It’s trying to win the “automate my back office” contest.
The Part Google Would Rather You Skip
Here’s where the catch comes in. Argon trails on frontier software engineering and science benchmarks against GPT-6 Astra, OpenAI’s current flagship. In plainer terms, if you’re a research lab trying to push the actual edge of what AI can discover or build from scratch, Astra reportedly still has the edge. Google’s own framing acknowledges this directly, with the company noting that “the frontier race remains close” even while touting the highest number of top scores across the test suite. That’s a carefully worded sentence, and it’s worth sitting with. Winning thirteen of eighteen benchmarks is impressive. Losing the two or three that measure raw frontier capability is the kind of detail that shows up in a competitor’s sales deck within a week.
Compared to Anthropic’s Claude Opus 5.5, the dynamic flips again. Argon is reportedly stronger on enterprise workflow automation, while Claude holds an edge on terminal-based tasks and anything requiring heavy post-training customization. Three labs, three different specialties. Nobody has lapped the field.
The Context Window Jump Nobody’s Talking About Enough
Buried under the benchmark chatter is a genuinely significant spec bump: Argon supports up to one million output tokens, a massive leap from the 64,000-token ceiling on Google’s prior release. Output tokens, for anyone who hasn’t had to think about this, are roughly what the model generates in response, not what you feed it. A sixteen-times jump in that ceiling means Argon can produce entire codebases, full-length technical reports, or multi-document legal filings in a single pass without the kind of chunking and stitching that currently eats up engineering time at companies building on top of these models.
Pricing That’s Designed to Undercut, For Now
Google is leaning hard into price as a weapon during the introductory period. Input tokens run $2 per million, output tokens $10 per million, and cached input tokens a startlingly cheap $0.10 per million — a 95 percent discount over standard rates. Run the math and that introductory pricing lands at roughly one-fifth the cost of GPT-6 Astra and about half of what Claude Opus 5.5 charges at standard rates. After the introductory window closes, prices roughly double, settling at $4 per million input tokens and $20 per million output tokens. Still competitive, but the aggressive undercut is clearly a land-grab move aimed at getting enterprise customers to build their pipelines around Argon before the honeymoon pricing ends.
Why You Can’t Actually Use It Yet
This is the part that undercuts the triumphant framing. Argon isn’t broadly available. Google is rolling it out first to what it calls “trusted cyber defenders” through its Fairwind Program, a gated group getting early access specifically for cybersecurity operations, while the model also undergoes a U.S. government pre-release review. Wider access for paid API customers and Google AI Ultra subscribers is coming “as soon as possible,” which is corporate speak for “we don’t have a firm date and don’t want to commit to one.”
That government review detail matters more than it might seem at first glance. Frontier labs have increasingly been looping in federal reviewers before wide releases, a pattern that’s become standard practice over the past year as regulators have grown more vocal about AI capabilities tied to cybersecurity and critical infrastructure. Google framing Argon’s first users as “cyber defenders” rather than developers or consumers is a signal about where the company sees the near-term value, and the risk, concentrated.
Google Is Already Eating Its Own Cooking
Internally, Google says its engineers are already running Argon against real infrastructure problems. The company reports a 40 percent improvement over baseline on quantum optimization tasks, along with projected data center memory savings in the 300 to 1,000 tebibyte range, and active use in migrating legacy C++ and Rust code, including work inside core kernel systems. Those aren’t demo-day flexes. Kernel-level code migration is tedious, high-stakes work that companies have historically been reluctant to hand to anything less than their most senior engineers. If Argon is actually doing meaningful work there, that’s a more convincing proof point than any benchmark chart.
What This Means
Argon doesn’t settle the AI model wars. It complicates them further. Google now has a legitimate claim to benchmark leadership on the metrics that matter most to enterprise buyers, but it’s conceding ground on the metrics that matter most to researchers chasing frontier capability. OpenAI and Anthropic both have obvious counters available, and given the pace of the last two years, expect one of them within a month, not a quarter.
For businesses actually shopping for a model to build on, the real story here isn’t who’s “winning.” It’s that the three leading labs have started to specialize, with Google chasing enterprise automation and legal-and-financial workflows, OpenAI defending its edge in raw frontier research, and Anthropic carving out territory in developer tooling and fine-grained customization. That’s a healthier market than a single model dominating every category, even if it makes for a less satisfying headline. The introductory pricing window is the thing to watch next. Whoever blinks first on cost after the honeymoon period ends will tell you a lot about how confident each lab actually is in its product.




