Grok 4.6 at $2/$6 vs GPT-5.6 Sol at $5/$30: the 2026 agent cost table
Summary. SpaceXAI released Grok 4.6 on 12 August 2026 at $2 per 1M input tokens and $6 per 1M output tokens, unchanged from Grok 4.5, and it scores 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol and behind Claude Opus 5 at 63. GPT-5.6 Sol lists at $5 and $30. On headline rates that is a 60% input saving and an 80% output saving. Three billing rules erase most of it for the wrong workload: Grok 4.6 bills every token in a request at the long-context rate of $4 and $12 once the prompt reaches 200,000 tokens, it has no Batch API discount, and its Priority tier costs 2x. GPT-5.6 Sol's batch rate is $2.50 and $15, which is half its standard rate. Separately, Anthropic confirmed that the Claude Sonnet 5 increase to $3/$15 scheduled for 1 September 2026 will not happen; $2/$10 is now the standard price.
The benchmark picture is also less flattering than the headline. On xAI's own published evaluations, Grok 4.6 scores 65.9% on DeepSWE v1.1 against 73% for GPT-5.6 Sol Max, and 26% on Terminal-Bench v3.0 against 34.6%. Buy it for agentic knowledge work, not for terminal-heavy software engineering.
What actually shipped on 12 August 2026
Grok 4.6 keeps the 500,000-token context window of Grok 4.5 and targets long-running agents: research across many steps, working through a codebase, turning a product idea into a first working version. xAI states that on longer trajectories the model began self-testing and verifying its own work before moving on. It is available in Cursor, Grok Build, the xAI API, and through OpenRouter, Vercel and Cloudflare.
Holding price flat across a model generation is the unusual part. Artificial Analysis measured a 5-point Intelligence Index gain over Grok 4.5 at unchanged pricing, and put the model on the intelligence-versus-cost Pareto frontier at a measured $0.84 per task.
The published benchmark spread
Two independent scorecards exist, and they disagree in a way worth understanding before you commit an agent fleet to a model.
xAI's announcement publishes a head-to-head table against Grok 4.5 High, GPT-5.6 Sol Max and Claude Fable 5 Max. Artificial Analysis publishes its own Intelligence Index and private agentic benchmarks. Where the two overlap, the ranking is consistent. Where they use different benchmark versions, the numbers diverge sharply: xAI reports Terminal-Bench v3.0 at 26% for Grok 4.6, while Artificial Analysis reports Terminal-Bench v2.1 at 88.4%. Those are different tests, not a contradiction, and quoting either without the version number is how vendor comparisons become fiction.
| Benchmark (source) | Grok 4.6 | GPT-5.6 Sol | Claude Fable 5 / Opus 5 |
|---|---|---|---|
| AA Intelligence Index (Artificial Analysis) | 61 | 61 | 62 Fable 5, 63 Opus 5 |
| GDPval-AA v2 Elo (xAI table) | 1753 | 1728 | 1741 Fable 5 Max |
| DeepSWE v1.1 (xAI table) | 65.9% | 73% | 70% Fable 5 Max |
| Terminal-Bench v3.0 (xAI table) | 26% | 34.6% | 34.1% Fable 5 Max |
| CursorBench v3.2 (xAI table) | 69.9% | 67.2% | 70.5% Fable 5 Max |
| APEX-Agents (xAI table) | 57.5% | 56.7% | 59.2% Fable 5 Max |
| Harvey LAB, Vals (xAI table) | 15.8% | 2.5% | 11.3% Fable 5 Max |
Read the table as a shape rather than a ranking. Grok 4.6 leads on legal and knowledge-work evaluations and on GDPval-AA, and trails clearly on the two software-engineering benchmarks that most coding-agent teams care about. The Harvey LAB gap is the widest single result in either direction.
The three billing rules that decide your bill
Headline rates are the smallest part of an agent's cost model. These three mechanics matter more.
Long-context billing is a cliff, not a slope. xAI's pricing documentation states that models with long-context pricing bill the long-context rates for all tokens in a request once its prompt reaches the model's long-context threshold. For Grok 4.6 that threshold is 200,000 tokens, and crossing it moves the whole request to $4 input, $1 cached input and $12 output. A 210,000-token agent turn costs double a 190,000-token turn on every token, not just on the excess. OpenAI applies a similar structure for the GPT-5.6 family, with long-context input at 2x and output at 1.5x. Anthropic does not: its documentation states that Claude 4.6 and later models include the full 1M-token context window at standard pricing, with a 900k-token request billed at the same per-token rate as a 9k-token request.
Batch changes the ranking. OpenAI's Batch and Flex tiers price GPT-5.6 Sol at $2.50 input and $15 output, half the standard rate. Anthropic's Batch API is also a 50% discount, putting Claude Opus 5 at $2.50 and $12.50 and Claude Sonnet 5 at $1 and $5. xAI publishes batch discounts only for grok-4.3 and the grok-4.20 family at 20%; Grok 4.6 has none. For any workload you can run asynchronously, Grok 4.6's input advantage over batched Sol shrinks to $2.00 against $2.50.
Speed tiers multiply everything. xAI's Priority Processing bills at 2x standard rates across input, output, cached and reasoning tokens. OpenAI's Fast mode (renamed from Priority on 30 July 2026) puts Sol at $10 and $60. Anthropic's Fast mode for Opus 5 is $10 and $50 and is not available with the Batch API. None of these appear in a headline comparison, and all of them are what an impatient product team turns on.
| Cost dimension | Grok 4.6 | GPT-5.6 Sol | Claude Opus 5 |
|---|---|---|---|
| Standard input / output per 1M | $2.00 / $6.00 | $5.00 / $30.00 | $5.00 / $25.00 |
| Cached input read | $0.50 | $0.50 | $0.50 |
| Long-context rates | $4.00 / $12.00 for the whole request at 200k+ | $10.00 / $45.00 long context | Full 1M window at standard pricing |
| Batch tier | None published | $2.50 / $15.00 | $2.50 / $12.50 |
| Fast or priority tier | 2x standard rates | $10.00 / $60.00 | $10.00 / $50.00 |
| Hosted web search | $5 per 1k calls | $10 per 1k calls plus content tokens | $10 per 1k searches plus content tokens |
| Context window | 500k | Long-context tier above 272k | 1M at standard pricing |
Turn efficiency beats per-token price
The most useful number in the Grok 4.6 coverage is not a price. Artificial Analysis measured, on its private AA-Briefcase benchmark of long-horizon agentic knowledge work, that Grok 4.6 resolves tasks in roughly 53 turns and 0.5 billion input tokens on average, against roughly 103 turns and 2.0 billion input tokens for Claude Opus 5 at maximum effort. Grok 4.6 scores 1577 Elo on that benchmark, behind the Claude Opus 5 family.
Artificial Analysis put the consequence plainly:
"Long-horizon agentic work accumulates context rapidly, so a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage well beyond its per-token pricing."
That is the correct way to price an agent. Micah Hill-Smith and George Cameron, co-founders of Artificial Analysis, have made the same argument publicly: the cost of a fixed level of intelligence keeps falling while the cost of completing a complex agentic task keeps rising, because long tasks repeatedly pass the same instructions, tool outputs and prior work back into the model. Token price is the multiplier. Turn count and token volume are the quantity, and the quantity varies by more than the price does.
Run the arithmetic on the AA figures and the gap is stark. At Grok 4.6's short-context input rate, 0.5 billion input tokens costs $1,000. At Opus 5's input rate, 2.0 billion input tokens costs $10,000 before caching. Prompt caching narrows that considerably, since a cache read on either platform costs $0.50 per 1M tokens, but the direction does not change.
Where each model is the right answer
Grok 4.6 is the pick for long-horizon agentic knowledge work, research agents, multi-turn customer service with tool use, and legal or document-heavy analysis, where its GDPval-AA and Harvey LAB results are strongest and its turn efficiency compounds. Keep individual requests under 200,000 tokens, and design context compaction into the harness rather than relying on the 500k window.
GPT-5.6 Sol wins on terminal-heavy and SWE-benchmark-shaped coding agents, where its 73% DeepSWE v1.1 and 34.6% Terminal-Bench v3.0 results lead this group, and for any pipeline you can move to Batch or Flex, where the 50% discount closes most of the price gap.
Claude Opus 5 is the answer where long context is unavoidable. It is the only one of the three that does not charge a long-context premium, so a workload that routinely runs 300k-token prompts can be cheaper on Opus 5 than on a nominally cheaper model that doubles its rate at the threshold. Watch the tokenizer: Anthropic documents that Claude 4.7 and later models use a newer tokenizer producing approximately 30% more tokens for the same text, which is a real cost multiplier that per-token comparisons hide.
Claude Sonnet 5 covers the volume tier. At $2 and $10 standard, $1 and $5 batched, it is the cheapest credible frontier-adjacent option in this comparison, and the September increase that teams had been planning migrations around has been cancelled.
The September reset that is not happening
Several teams built Q3 migration plans around Claude Sonnet 5's introductory pricing expiring on 31 August 2026. Anthropic's pricing documentation now carries an explicit note: the $2/$10 pricing announced at launch as introductory through 31 August 2026 is the standard price, and the previously scheduled increase to $3/$15 on 1 September 2026 will not occur.
If you queued a migration for late August on that basis, the business case has changed. Our earlier analysis of the Claude Sonnet 5 migration and September price cliff should be read alongside this correction. The tokenizer point in that piece still stands; the price cliff does not.
This is the second time in six weeks that a scheduled frontier price move has been reversed or cut. OpenAI cut Terra and Luna rates on 30 July 2026 and renamed Priority to Fast mode the same day. The practical lesson for anyone modelling 2027 spend: pin your comparison to a dated snapshot of the vendor's own pricing page, and re-run it before you sign anything.
Hidden line items your model comparison probably missed
Hosted tools are billed separately and the rates differ by more than the token rates do. xAI charges $5 per 1,000 calls for web search, X search and code execution, $10 per 1,000 for file attachment search, and $2.50 per 1,000 for collections search. OpenAI charges $10 per 1,000 web search calls plus the search content tokens at model rates, and bills Hosted Shell and Code Interpreter containers per 20-minute session by memory size. Anthropic charges $10 per 1,000 web searches, makes web fetch free beyond token costs, and gives each organisation 1,550 free hours of code execution per month before charging $0.05 per container-hour.
Data residency is a multiplier, not a feature flag. Anthropic applies a 1.1x multiplier across every token category when inference_geo is pinned to "us" on Claude 4.6 and later. OpenAI applies a 10% uplift on regional processing endpoints for models released on or after 5 March 2026. For an Indian buyer routing through a specific region for contractual reasons, that 10% sits on top of everything else in the table.
Cache write pricing is asymmetric. Anthropic charges 1.25x base input for a 5-minute cache write and 2x for a 1-hour write, with reads at 0.1x. OpenAI lists a separate cache-write rate of $6.25 for Sol. xAI publishes a cached-input rate of $0.50 for Grok 4.6, up from $0.30 on Grok 4.5, so the cache economics got slightly worse in this generation even as headline pricing held.
xAI also charges a $0.05 usage guideline violation fee per request for violations caught before generation in the Responses API. It is small, but it is a line item that no competitor has, and a misconfigured agent loop can produce a lot of them.
India-specific considerations
All three vendors bill in USD, so an Indian buyer carries the exchange-rate exposure on a cost line that is already volatile month to month. That argues for the same discipline we recommend on cloud spend: instrument per-feature token attribution before you negotiate, not after.
Data residency deserves particular attention under the DPDP Act. If prompts carry personal data and your contractual position requires regional processing, both Anthropic's 1.1x inference_geo multiplier and OpenAI's 10% regional uplift apply on top of the rates above, and they apply to every token category including cache reads. Model that before you pick on headline price. Our guide to data residency and DPDP cloud architecture covers the architectural side.
For teams standardising a model portfolio rather than a single model, see our frontier model comparison for 2026 and the tier-selection breakdown in GPT-5.6 Sol, Terra and Luna tier selection.
How to run this comparison on your own workload
Instrument turn count and input tokens per completed task, not cost per million tokens. The AA-Briefcase figures show a 4x spread in input tokens between two models scoring within two points of each other. No per-token table can tell you that.
Measure your prompt-size distribution against the thresholds. If more than a small fraction of requests cross 200,000 tokens, Grok 4.6's effective rate is not $2 and $6. Plot the histogram before you choose.
Separate the workload into interactive and asynchronous. Everything asynchronous should be priced at batch rates, which changes the ranking. Everything interactive should be priced with the speed tier you will actually enable under load.
Add hosted tool calls to the model. An agent that runs six searches per task at $10 per 1,000 calls adds $0.06 per task, which is a material fraction of a measured $0.84 per task.
Re-run the comparison monthly. Two scheduled price moves have already been reversed or cut in the last six weeks.
FAQ
Is Grok 4.6 actually as good as GPT-5.6 Sol?
They tie at 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks. The tie hides a split: xAI's own table shows Grok 4.6 ahead on GDPval-AA v2 and Harvey LAB, and behind on DeepSWE v1.1 at 65.9% against 73% and Terminal-Bench v3.0 at 26% against 34.6%.
What is the catch in Grok 4.6's $2 and $6 pricing?
Three things. Once a prompt reaches 200,000 tokens, xAI bills the entire request at $4 input and $12 output, not just the excess. There is no Batch API discount for Grok 4.6. And Priority Processing bills at 2x standard rates across input, output, cached and reasoning tokens.
Does Claude Opus 5 charge extra for long context?
No. Anthropic's pricing documentation states that Claude 4.6 and later models include the full 1M-token context window at standard pricing, and that a 900k-token request is billed at the same per-token rate as a 9k-token request. Prompt caching and batch discounts apply at standard rates across the full window.
Did Claude Sonnet 5 get more expensive on 1 September 2026?
No. Anthropic's pricing page states that the $2 and $10 per million input and output token pricing, announced at launch as introductory through 31 August 2026, is now the standard price, and that the previously scheduled increase to $3 and $15 per million tokens on 1 September 2026 will not occur.
Which model is cheapest for a long-running research agent?
It depends on turn efficiency, not headline price. Artificial Analysis measured Grok 4.6 completing AA-Briefcase tasks in roughly 53 turns and 0.5 billion input tokens against roughly 103 turns and 2.0 billion for Claude Opus 5 at maximum effort, which is a larger spread than the per-token price difference.
How much do hosted tools add?
xAI charges $5 per 1,000 web search, X search and code execution calls, and $10 per 1,000 file attachment searches. OpenAI charges $10 per 1,000 web search calls plus search content tokens at model rates. Anthropic charges $10 per 1,000 web searches and provides web fetch at no additional cost beyond tokens.
Does data residency change the price?
Yes, on two of the three platforms. Anthropic applies a 1.1x multiplier on all token categories when inference geography is pinned to the United States for Claude 4.6 and later. OpenAI charges a 10% uplift on regional processing endpoints for models released on or after 5 March 2026, which stacks with everything else.
Why do two sources report such different Terminal-Bench scores?
Because they run different versions. xAI's announcement reports Terminal-Bench v3.0 and gives Grok 4.6 26%. Artificial Analysis reports Terminal-Bench v2.1 and gives it 88.4%. Both are accurate for the test they ran, which is why a benchmark figure without its version number is not usable for procurement.
How eCorpIT can help
eCorpIT builds and operates production AI agents for teams in India and abroad, and the model choice is usually the last decision we make, not the first. We instrument turn count, token volume and prompt-size distribution on your actual workload, price it against current vendor rates including batch, speed and residency multipliers, and design the harness so context compaction keeps requests under the thresholds that trigger long-context billing. We are ISO 27001:2022 certified, CMMI Level 5 appraised and MSME certified. To model your agent spend against real usage data, contact us.
References
- Introducing Grok 4.6 - SpaceXAI, 12 August 2026
- xAI API pricing - SpaceXAI Docs
- Grok 4.6 model page - SpaceXAI Docs
- Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency - Artificial Analysis, 12 August 2026
- Artificial Analysis model page for Grok 4.6 - Artificial Analysis
- OpenAI API pricing - OpenAI
- GPT-5.6 Sol model page - OpenAI
- Advancing the price-performance frontier with GPT-5.6 - OpenAI
- Claude pricing - Anthropic
- Claude Opus - Anthropic
- Claude Sonnet - Anthropic
- Artificial Analysis on the cost of intelligence for agentic tasks - Micah Hill-Smith and George Cameron, Artificial Analysis
- Grok 4.6 is out, undercutting AI prices of rivals - AI Business
- Grok 4.6 released: benchmarks, pricing, and what it means for agent builders - DEV Community
- Introducing Grok 4.5 - SpaceXAI
Last updated: 14 August 2026.
Top comments (0)