Grok 4.6: Release Date, Pricing, Benchmarks, and What's Actually New
Grok 4.6 is here. See the real API pricing, benchmark scores vs GPT-5.6 and Claude Opus 5, context window limits, and whether it's worth switching to.
Long Nguyen
Fullstack Developer · AI Engineer · Researcher
xAI (now operating under its parent, SpaceXAI) shipped Grok 4.6 on August 12, 2026, and the headline pitch is simple: a meaningful step up from Grok 4.5 in coding and agentic performance, at the same API price. Whether that pitch holds up depends on which number you look at — so here's the breakdown, without the launch-day hype.
What Is Grok 4.6?
Grok 4.6 is xAI's latest frontier language model, built as a refinement of the same 1.5-trillion-parameter foundation that powered Grok 4.5, rather than a larger base model. The gains come almost entirely from post-training — upgraded supervised fine-tuning and reinforcement learning, including xAI's "Grok Build" coding harness, which trains the model against real coding tasks.
That's a notable strategic choice. Instead of scaling parameters, xAI is testing how far post-training alone can push a model — and early scores suggest the answer is "quite far."
The model supports text and image input with text-only output, a 500,000-token context window, function calling, structured outputs, and configurable reasoning effort. It's available now through the xAI API, Grok Build, Cursor, Grok Bot, OpenRouter, Vercel, and Cloudflare.
Grok 4.6 Benchmarks: How It Actually Ranks
On the Artificial Analysis Intelligence Index — one of the more widely cited independent benchmark aggregates — Grok 4.6 scores 61, putting it roughly level with GPT-5.6 Sol and about five points above Grok 4.5. That places it just behind Anthropic's Claude Opus 5 and Fable 5, which currently hold the top two spots on the same index.
A few specifics worth knowing:
- Grok 4.6 leads on CursorBench, FrontierCode, and AA-Briefcase, but still trails GPT-5.6 Sol on DeepSWE and a couple of other coding-specific evaluations.
- On agentic task completion, Grok 4.6 reportedly finishes complex multi-step workflows in around 53 steps, compared to roughly 103 for Claude Opus 5 — at a significantly lower price per task.
- Some early testers report an ELO score near 1753 in head-to-head model arenas, though independent confirmation from LMSYS Chatbot Arena and Artificial Analysis was still pending shortly after launch.
The honest caveat: these are largely vendor-reported or self-selected benchmark results. xAI states its third-party comparisons use the best publicly available scores for competing models, which means testing conditions, reasoning settings, and tool access may not be identical across the board. Treat the numbers as directionally useful, not as a settled final ranking.
Grok 4.6 Pricing: The Part Most Comparisons Get Wrong
The headline API rate is $2 per million input tokens and $6 per million output tokens — unchanged from Grok 4.5, and roughly 60% cheaper than Claude Opus 5 ($5/$25) or GPT-5.6 Sol ($5/$30) at list price.
But that $2/$6 figure only applies below a 200,000-token prompt. Once a request crosses that threshold, xAI's long-context pricing kicks in and doubles the rate — to $4 per million input tokens and $12 per million output tokens — and importantly, the higher rate applies to every token in that request, not just the tokens past the threshold.
| Standard (under 200K tokens) | Long-context (200K+ tokens) | |
|---|---|---|
| Input | $2 / 1M tokens | $4 / 1M tokens |
| Cached input | $0.50 / 1M tokens | $1 / 1M tokens |
| Output | $6 / 1M tokens | $12 / 1M tokens |
Cached input pricing also quietly rose from $0.30 to $0.50 per million tokens on the standard tier compared to Grok 4.5 — a detail that didn't make it into the launch announcement. If your workload leans on long documents or large repos, model the long-context tier before assuming the "half the price of frontier models" framing applies to you.
For consumer access, Grok remains available through X/Grok Bot on tiered plans, with the top-end SuperGrok Heavy plan priced around $300/month.
Grok 4.6 vs GPT-5.6 Sol vs Claude Opus 5
Here's how the three currently stack up on publicly available data:
| Model | AA Intelligence Index | Input / Output Price (per 1M tokens) | Context Window |
|---|---|---|---|
| Grok 4.6 | 61 | $2 / $6 | 500K |
| GPT-5.6 Sol | 61 | $5 / $30 | — |
| Claude Opus 5 | ~63–65 (ranked #1) | $5 / $25 | — |
Grok 4.6 ties GPT-5.6 Sol on raw intelligence scoring while undercutting it substantially on price, and it edges out the open-weight Kimi K3 model as well. Claude Opus 5 and Fable 5 still lead the overall index, so "best model" and "best value" aren't currently the same answer — which one matters depends on whether your workload is cost-sensitive or ceiling-sensitive.
Is Grok 4.6 Worth Switching To?
A few practical takeaways if you're deciding whether to move a workload onto Grok 4.6:
- Coding and agentic workflows are where the upgrade is most visible — xAI's Grok Build harness was trained specifically against real coding tasks, and the step-count efficiency on agentic benchmarks is a real, measurable difference.
- Cost-sensitive, high-volume use cases benefit the most from the price gap, as long as you're staying under the 200K-token long-context threshold.
- Long-document or large-codebase workloads need to price out the long-context tier separately — the doubled rate applies retroactively to the whole request, which changes the math fast.
- Enterprise and regulated deployments should weigh vendor history alongside benchmark scores; procurement teams evaluating any model typically look past raw capability numbers to governance and reliability track record.
What's Next: Grok 4.7
xAI has already signaled that a larger, 2.1-trillion-parameter model — Grok 4.7 — is expected within weeks of the 4.6 release, with Grok 5 targeted before the end of 2026. If that cadence holds, Grok 4.6 may end up being a relatively short-lived release rather than xAI's flagship for long. Worth keeping in mind if you're planning infrastructure around a specific model version rather than an API endpoint that xAI will keep upgrading underneath you.
FAQ
Frequently asked questions
When was Grok 4.6 released?
August 12, 2026, across the xAI API, Grok Build, Cursor, and Grok Bot.
What is the Grok 4.6 API price?
$2 per million input tokens and $6 per million output tokens for prompts under 200,000 tokens. Above that threshold, pricing doubles to $4/$12 per million tokens for the entire request.
What is Grok 4.6's context window?
500,000 tokens, with long-context pricing applying above 200,000 tokens.
Is Grok 4.6 better than GPT-5.6 Sol?
They're roughly tied on the Artificial Analysis Intelligence Index (both around 61), with Grok 4.6 priced significantly lower and stronger on agentic step-efficiency, while GPT-5.6 Sol leads on some coding-specific benchmarks like DeepSWE.
Is Grok 4.6 better than Claude Opus 5?
No — Claude Opus 5 currently ranks higher on the Intelligence Index, but Grok 4.6 is considerably cheaper per token, which shifts the calculus for cost-sensitive or high-volume workloads.