Grok 4.6 Sells Frontier Intelligence at $2/M Tokens

Grok 4.6 Sells Frontier Intelligence at $2/M Tokens

SpaceXAI shipped Grok 4.6 on August 12, 2026.

And put a new floor under the frontier price war: $2 per million input tokens, $6 per million output tokens.

The model is SpaceXAI’s frontier release for coding, agentic tasks. And knowledge work, live on the xAI API as model id grok-4.6 and in Cursor, Grok Build.

And the Grok Bot app, with OpenRouter, Vercel, and Cloudflare distributing it as partners.

One report scores it at 61 on the Artificial Analysis Intelligence Index, which SpaceXAI describes as a composite of nine benchmarks, matching GPT-5.6 Sol at its maximum reasoning level.

That combination, frontier-adjacent scores at mid-tier prices, is why this launch matters if you pay for tokens. When capability ties at the top, the decision stops being “which model is smartest” and becomes “what my workload actually costs on each one.” That second question hides a twist in Grok 4.6’s fine print. And I’ll get to it.

The Price That Started the Fight

The launch numbers are $2 per million input and $6 per million output, per SpaceXAI’s announcement.

The release notes add a cached-input rate of $0.50 per million tokens below 200k prompt tokens. And that number matters more than it first appears. Agents re-send the same system prompt, tool definitions. And file context on every step of a run. So cache pricing decides whether a long-running agent stays affordable past week two.

Two more numbers complete the picture.

A fast variant ships at twice the price, which turns raw speed into a line item you either pay for or skip.

And Cursor and Grok Build users get 2x included usage for the first week, which is the cheapest real-workload evaluation window you will get on this model.

Read the price war correctly. This is not a lab cutting prices out of generosity; it is a lab pricing for the volume it wants to capture. Every builder who flips a default endpoint because of a headline number is that volume.

A 61 on the Index, With an Asterisk

SpaceXAI’s claim is that Grok 4.6 “achieves frontier intelligence across several agentic coding and knowledge work benchmarks” and matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index. One report puts the actual score at 61, matching GPT-5.6 Sol at its maximum reasoning level and trailing Fable 5 Max by one point.

Treat those numbers with the right amount of respect.

A tie on a composite of nine benchmarks tells you two models are close on average across nine tasks.

And “close on average” is not what you ship on. You do not run composites; you run document extraction, codebase edits. And support triage, and a one-point index gap says nothing about which model wins those specific jobs.

The asterisk that matters: the tie is with Sol at its maximum reasoning level.

That framing cuts both ways.

It says Grok 4.6 reached parity with a rival running full effort. And it reminds you that reasoning effort is a dial you control and pay for on every single call.

The 200k Token Cliff Buried in the Fine Print

The pricing docs carry a second price list, and it is the most important part of this launch. Below 200k prompt tokens you pay $2 input / $0.50 cached / $6 output per million. Above 200k prompt tokens, the whole request bills at $4 input / $1 cached / $12 output. The doubled rate applies to the entire request, not just the tokens past the threshold.

Now line that up against what the model is for. Grok 4.6 has a 500k context window. And 9to5Mac’s coverage describes it as designed for long-running agents, coding, knowledge work, and more ambitious visual projects. Long-running agents are precisely the workloads that accumulate context past 200k tokens, as every step appends tool output and results onto the prompt.

The headline price and the flagship feature pull against each other. The $2/$6 rate lives in the small-context world. And the 500k window invites you into the world that bills at $4/$12. An agent that reads a codebase and iterates lives in the expensive tier. And the launch-page math stops applying somewhere in hour one of its first real run.

One more cost lever: Grok 4.6 is a reasoning model with extended chain-of-thought.

And it adds a new xhigh reasoning-effort setting above the ladder Grok 4.5 shipped with, per Orca Router’s writeup.

Reasoning burns output tokens, output is the expensive side of both price tiers. And the release notes specify no text output limit.

A model that thinks longer at xhigh is a model that bills longer.

What I’d Actually Do This Week

The first-week window is the cheapest testing you will get, so spend it deliberately:

– Point one real workload at model id grok-4.6 through OpenRouter, Vercel, or Cloudflare, whichever you already use, so you test without new billing plumbing.
– Claim the 2x included usage in Cursor and Grok Build before the week closes, and spend it on tasks where you already know what good output looks like.
– Run the same task at default reasoning effort and at xhigh, then compare token counts, not vibes. The delta is a real price tag.
– Audit your prompt sizes. Your effective price is $2/$6 or $4/$12 depending on which side of 200k your traffic lives, and most people guessing this number guess wrong.

That last bullet is the whole post.

The price war’s real weapon is not the $2 headline; it is knowing which tier your workload lives in, since the same model bills two very different rates depending on how much context you feed it.

Frontier capability at commodity input prices changes the buying logic for small operators.

You no longer need to defend paying frontier rates for agent work. But you do need to defend your context size. Pull last month’s token logs, price your traffic at both tiers. And make the switch call on your own numbers instead of launch-day slides. The audit takes an hour. And it is the difference between winning this price war and discovering you were never in the cheap tier to begin with.

Leave a Reply

Your email address will not be published. Required fields are marked *