
Gemini 4 Argon can emit 1 million tokens in a single run. And every headline number in its launch deck was produced by the company that built the model. Google announced Argon on September 30, 2026 as its new frontier model, aimed at real-world software engineering, company knowledge work like legal and finance, and cybersecurity defense. It ships with a 1M-token output limit, up from 64K on previous Gemini models, at an introductory price of $2 per million input tokens and $10 per million output tokens.
Right now it sits with vetted cybersecurity defenders through Google’s Fairwind Program, with paid API customers and Google AI Ultra subscribers next in line and no public date for that as of October 3, 2026.
That’s the whole announcement. The rest of this post is about which parts you should trust. And which parts should change how you price your own AI workflows before the model ever reaches your API key.
Google Ran the Benchmarks and Won Them
Google reported that Argon sets a new state of the art on DeepSWE v1.1 at 77.9%, a benchmark measuring real-world long-horizon software engineering tasks. Google also reported 91.7% on LVBench, state of the art for long video understanding.
Both numbers come from Google’s own launch post, on benchmarks Google chose to headline, with the model’s maker as the scorer.
SiliconANGLE captured the framing exactly: Google claims Argon beats rival models from Anthropic and OpenAI on most of Google’s benchmarks. Read that sentence again. The vendor picked the tests, ran the tests, and reported the results. That’s not lying, it’s selection, and selection is the oldest trick in launch marketing.
The most honest number in the coverage is the one Google didn’t win outright.
On CWE-bench v1, a software-vulnerability-remediation test, Argon tied GPT-6 Astra at 68%, per SiliconANGLE’s reporting. That tie is the figure I’d anchor any buying decision on. Because it’s the only externally reported comparison in the launch coverage and it came out level. When the self-graded exam and the proctored exam disagree, trust the proctored one.
The 1M-Token Ceiling Is a Cost Decision Disguised as a Feature
Google said the expansion from 64K to 1M output tokens exists so the model can “think deeply and generate hundreds of thousands of tokens in a single trajectory” and “solve tough problems in one go.” That capability claim has an invoice attached.
Output is the expensive direction: $10 per million output tokens versus $2 for input.
Run the math on the ceiling itself.
At $10 per million output tokens, a run that maxed the old 64K limit spent at most $0.64 on output.
A run that maxes the new 1M limit spends $10.00 on output. Nothing about your task changed.
But the worst case per single call moved by more than ten times, given that the ceiling moved from 64K to 1M and your worst case moves with the ceiling.
There’s a discount worth knowing about too. Google priced cached input at 95% off the input token price. And agent loops that re-send the same context on every step are exactly what that pricing rewards. I price my own automated workflows by output tokens as that’s where loops actually spend money. And a model given permission to think for hundreds of thousands of tokens converts “smarter” into “more expensive per decision” faster than any capability gain rescues it. Before you celebrate the headroom, find out what your tasks do with it.
Defenders Got It First Since the Capability Cuts Both Ways
The rollout order is the real product message. Argon went first to trusted cyber defenders through the Fairwind Program. And Google’s announcement contains a capability claim worth reading slowly: “Argon can autonomously find, validate. And patch critical software vulnerabilities.” A model that finds vulnerabilities finds them for whoever is holding it. Google said it will “iterate on guardrails” using early-tester feedback before making Argon available to developers, enterprises.
And consumers, which is the company telling you in plain language that the safety work is not finished.
For a solo operator or a small team, the defense framing is not the interesting part.
Finding and patching vulnerabilities autonomously is exactly the maintenance work that eats lean operations alive, the dependency updates and remediation write-ups that never make the roadmap given that shipping does. If the claim survives contact with independent testing, the move is preparing those workflows now. So the day access opens you’re pointing an existing pipeline at a new model instead of designing one from scratch under deadline.
What To Do Before You Have Access
– Don’t block your roadmap on it. No general-release date existed in reporting as of October 3, 2026, and the first wave goes to Google AI Ultra subscribers and paid API customers, not the general API.
– Reprice your workflows at $2 and $10 per million now. Anything you costed at other rates needs the model rerun at these prices before Argon lands, not after.
– Cache aggressively. Cached input at 95% off the input price is built for loops that resend context, and most loops resend context constantly.
– Measure your output-token share per task this week. If your workflows generate long documents or long diffs, the 1M ceiling is where your bill will live.
– Treat 77.9% and 91.7% as Google’s homework pending independent grading. The 68% tie with GPT-6 Astra is the only externally reported comparison so far, and it’s level.
The Call
Gemini 4 Argon is a serious claim wrapped in self-graded evidence. The token ceiling, the cache discount. And the price are public and concrete, so build your cost models on those today. The benchmark wins are Google’s own numbers on Google’s chosen tests. So hold your architectural commitments until someone outside Mountain View runs them. Vendors announce frontier performance every quarter; the invoice and the tie score are the parts that still mean something next quarter.
Pull your API logs from last month and recompute two or three production workflows at $2 per million input and $10 per million output, with cached pricing applied to the repeated context.
Do it before Argon reaches your tier, as the operators who know their per-task cost at the new rates will adopt on day one with numbers.
And everyone else will adopt since the launch post sounded confident.
