Loading...


Updated 22 Aug 2026 • 3 mins read

AI Grok API pricing in 2026 is per million tokens, led by Grok 4.6 at $2 input and $6 output, with cached input at $0.50. This guide covers every current Grok model, the long-context surcharge that kicks in at 200K tokens, tool-call fees, and how to reduce your Grok bill.
Grok's pricing has an unusual shape, and the shape is the story. On output tokens it is cheap, its flagship charges three times its input rate where most rivals charge five or six, which makes it attractive for output-heavy work. But cross a single threshold, 200,000 tokens in a prompt, and the price of the entire request roughly doubles. Grok rewards teams that understand its two quirks and quietly punishes those who feed it giant prompts without checking the meter.
This guide lays out xAI Grok API pricing in 2026, model by model, with the caching discount and the long-context surcharge that decide what you actually pay, plus the levers to keep the bill in check. Note that the developer API is billed separately from consumer Grok subscriptions like SuperGrok.
The short version Grok's flagship, Grok 4.6, costs $2.00 per million input tokens and $6.00 per million output, with cached input at $0.50, for prompts under 200K tokens. Above 200K, the whole request jumps to $4/$12. Cheaper tiers include Grok 4.3 at $1.25/$2.50 and Grok 4.1 Fast at $0.20/$0.50 with a 2M-token window. Built-in web and X search tool calls bill around $5 per 1,000, and Priority Processing costs 2x standard. Model choice, caching, and staying under 200K are the biggest levers.
Like other LLM APIs, Grok bills per million tokens, input and output counted separately, with no monthly subscription, you pay for what you send and receive. Three things beyond the headline rate shape the bill: cached input (repeated prompt prefixes served at a steep discount), the 200K long-context surcharge, and separately billed tool calls when Grok invokes web, X, code, or file search. If per-token pricing is new to you, our token economics and TokenOps guide covers the fundamentals.
Here is the current xAI lineup and what each model costs per million tokens for prompts under 200K.
| Model | Input / 1M tokens | Output / 1M tokens | Cached input / 1M tokens | Context |
|---|---|---|---|---|
| Grok 4.6 (flagship) | $2.00 | $6.00 | $0.50 | 500K |
| Grok 4.5 | $2.00 | $6.00 | $0.30 | 500K |
| Grok 4.3 / 4.20 | $1.25 | $2.50 | $0.20 | 1M |
| Grok Build 0.1 (coding) | $1.00 | $2.00 | — | 256K |
| Grok 4.1 Fast | $0.20 | $0.50 | — | 2M |
A few things stand out. Grok 4.6 keeps a 3:1 output-to-input ratio, cheaper on output than most flagships. Grok 4.1 Fast pairs near-budget pricing with a 2M-token window, making it the value pick for classification, extraction, and long-document work. And Grok Build 0.1 is the dedicated, discounted coding model. All rates should be confirmed on the xAI console, since xAI has repriced models mid-quarter before.
This is the one that catches first-time budgets. Once a prompt reaches 200,000 tokens, xAI bills the entire request at the higher long-context tier, not just the tokens above 200K. On Grok 4.6 that means $4 per million input and $12 per million output for the whole call. A 210K-token prompt is billed as 210K at the high rate, not 200K low plus 10K high. If your workload flirts with the boundary, keeping prompts under it is a real saving.
Grok's built-in web and X search tools bill separately, around $5 per 1,000 calls, when the model chooses to invoke them, so an agent that searches on every turn adds a line the token rate hides. Priority Processing, a higher-scheduling-priority service tier, costs 2x standard token prices and applies only when the response confirms it ran at priority. Neither is free, so enable them deliberately.
These are the same levers that control any token bill, covered in our LLM cost optimization and FinOps for AI guides. For how Grok stacks up against other providers, see our Gemini API pricing and Anthropic vs OpenAI comparisons.
Grok's positioning is clearest at the edges of the lineup. At the top, Grok 4.6's cheap output makes it competitive for output-heavy generation among frontier models, as long as prompts stay under the 200K line. At the bottom, Grok 4.1 Fast is one of the strongest quality-per-dollar options anywhere, with a 2M-token window that rivals charge far more to match. The middle tiers, Grok 4.3 and Grok Build, cover general and coding work at mid-range rates. The right Grok is rarely the flagship by default; it is the cheapest tier that clears your quality bar for the task in front of it.
Grok is priced for people who read the meter. Its cheap output and deep caching discount reward output-heavy and repetitive workloads, while its 200K surcharge and per-call tool fees punish the ones that ignore structure. There is no single Grok price, only the price of the model and prompt shape you choose.
So treat model selection and prompt size as budget decisions, not defaults. Pick the smallest tier that does the job, cache what you send twice, keep prompts under the long-context line, and turn tools on with intent. Do that, and Grok is one of the better-value frontier APIs on the market. Ignore it, and a few oversized prompts will teach you the rate card the expensive way.
Grok's flagship Grok 4.6 costs $2.00 per million input tokens and $6.00 per million output, with cached input at $0.50, for prompts under 200K tokens. Cheaper tiers include Grok 4.3 at $1.25/$2.50 and Grok 4.1 Fast at $0.20/$0.50. The developer API is separate from consumer Grok subscriptions.
Once a prompt reaches 200,000 tokens, xAI bills the entire request at a higher tier, roughly double. On Grok 4.6 that is $4 per million input and $12 per million output for the whole call, not just the tokens above 200K, so a 210K-token prompt is billed entirely at the higher rate.
Grok 4.1 Fast at $0.20 per million input and $0.50 output is the cheapest, and it pairs a 2M-token context window with strong quality-per-dollar, making it the value pick for classification, extraction, and long-document work.
Yes. Grok's built-in web and X search tools bill separately, around $5 per 1,000 calls, when the model invokes them. Priority Processing, a higher-priority service tier, costs 2x standard token prices and applies only when a response runs at priority.
Repeated prompt prefixes such as a fixed system prompt or repository context are served from cache at a discount, $0.50 per million on Grok 4.6 versus $2.00 for fresh input, a 75 percent reduction. For agents and chat apps that resend context, cached input often dominates the bill.
On output tokens, Grok's flagship is competitive because it keeps a 3:1 output-to-input ratio where many rivals charge 5:1 or 6:1. Whether it is cheaper overall depends on your prompt shape, how much context you cache, and whether you stay under the 200K long-context line.
Use the cheapest model that clears your quality bar, cache repeated context, keep prompts under 200K tokens to avoid the surcharge, and gate tool calls and Priority Processing so they fire only when needed.