Loading...


Updated 21 Sep 2026 • 5 mins read

LLM token costs in 2026 span a 300-fold range, from about $0.035 per million input tokens on the cheapest models to $10 on frontier tiers. This guide lists current per-token rates for every major model, explains why output costs more than input, and covers the levers that reduce spend most.
Two teams build the same feature. One spends $400 a month running it, the other spends $40,000. Neither wrote better code than the other, and the difference is not traffic. It is model choice, prompt design, and whether anyone checked the output token count. Token pricing is the most consequential technical decision most AI teams never consciously make, because it is made by default, on the first day, when someone picks the model they read about most recently and never revisits it.
The spread is genuinely that wide. Across the models available in September 2026, input rates run from about $0.035 per million tokens at the bottom to $10 at the top, a factor of roughly 300, and output rates run from $0.14 to $50. A model that is 300 times more expensive is not 300 times better at classifying support tickets. This guide lays out what every major model actually charges, explains the mechanics that make bills larger than the rate card suggests, and covers the levers that reliably cut spend.
The short answer: LLM token costs in 2026 range from about $0.035 per million input tokens (Amazon Nova Micro) to $10 (Claude Fable 5.1 and GPT-6 Astra), with output typically priced at three to five times input. Mainstream production models cluster between $1 and $5 per million input tokens: Claude Sonnet 5 at $2 input and $10 output, Claude Haiku 4.5 at $1 and $5, Gemini 3.1 Pro at $2 and $12, and GPT-5.4 at $2.50 and $15. The cheapest way to cut an LLM bill is not negotiating rates but routing each request to the smallest model that clears your quality bar, then applying prompt caching and batch processing, which cut 90 percent and 50 percent respectively on eligible traffic.
Every major provider bills the same way: per token, split into input and output, quoted per million tokens. A token is a chunk of text, roughly three-quarters of a word in English, so a thousand tokens is about 750 words. Input tokens are everything you send, the system prompt, the conversation history, retrieved documents, and the user's message. Output tokens are everything the model generates. The mechanics are covered in depth in our token economics and TokenOps guide, but three properties of that model explain most surprise bills.
First, output costs more than input, usually three to five times more, because generating text is computationally harder than reading it. A verbose model is an expensive model, and capping maximum output length is one of the few changes that reduces cost without reducing quality.
Second, input accumulates. In a conversation, every turn resends the entire history, so the twentieth message in a chat may carry twenty times the input of the first. The same applies to retrieval-augmented systems that stuff large documents into context on every call.
Third, and most costly in 2026, reasoning tokens are billed as output. Models with extended thinking generate intermediate reasoning before their visible answer, and on most providers those tokens are charged at the output rate even though the user never sees them. A three-sentence answer can sit on top of several thousand billed tokens, which is why a page that models cost from visible output alone will understate the bill substantially.
The table below lists standard list rates per million tokens in US dollars, as of September 2026. Batch processing halves these on every provider that offers it, and caching reduces input further. Rates change frequently, so confirm on each provider's official pricing page before building a budget.
| Model | Input / 1M | Output / 1M | Provider | Positioning |
|---|---|---|---|---|
| Amazon Nova Micro | $0.035 | $0.14 | Amazon | Cheapest available; simple, high-volume tasks |
| Amazon Nova Lite | $0.06 | $0.24 | Amazon | Low-cost general use |
| GPT-5-nano | $0.05 | $0.40 | OpenAI | Classification, routing, tagging |
| DeepSeek V4 Flash | $0.22 | $0.66 | DeepSeek | Cheapest frontier-class (off-peak) |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Budget tier with 1M context | |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | High-throughput budget tier | |
| DeepSeek V4 Pro | $0.66 | $1.98 | DeepSeek | Low-cost heavy reasoning (off-peak) |
| Amazon Nova Pro | $0.80 | $3.20 | Amazon | Balanced Amazon tier |
| Claude Haiku 4.5 | $1.00 | $5.00 | Anthropic | Fast, low-cost Claude |
| Gemini 3.x Flash | $1.50 | $7.50 | Google's workhorse ($0.75/$3.75 intro) | |
| Claude Sonnet 5 | $2.00 | $10.00 | Anthropic | Default production Claude |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4/$18 above 200K tokens; 2M context | |
| xAI Grok 4.6 | $2.00 | $6.00 | xAI | Low output rate relative to input |
| GPT-5.4 | $2.50 | $15.00 | OpenAI | OpenAI production workhorse |
| Claude Opus 5 | $5.00 | $25.00 | Anthropic | Complex agentic and enterprise work |
| GPT-5.5 | $5.00 | $30.00 | OpenAI | Frontier reasoning |
| Claude Fable 5.1 | $10.00 | $50.00 | Anthropic | Highest Claude tier; $0.25/M cache reads |
| GPT-6 Astra | $10.00 | $50.00 | OpenAI | Frontier computer-use model |
The first pattern is that the middle is crowded and the extremes are not. Between $1 and $3 per million input tokens sits a dense cluster of genuinely capable models from four different providers, all of which handle the large majority of production workloads well. Competition has made the mid-tier a commodity, which is good news: switching between Sonnet 5, Gemini 3.1 Pro, GPT-5.4, and Grok 4.6 is a quality and latency decision more than a cost one.
The second is that the frontier tier costs exactly twice the flagship tier at both ends, $10 and $50 against $5 and $25. That is a deliberate pricing signal from both Anthropic and OpenAI: frontier models are positioned for work that genuinely fails on the tier below, not as a general upgrade. If a task succeeds on Opus 5 or GPT-5.4, running it on Fable 5.1 or Astra doubles the cost for nothing.
The third is that output-to-input ratios vary more than the headline rates suggest, and they change which model is cheapest for your shape of work. Grok 4.6 charges 3x output over input, most models charge 5x, and GPT-5.5 charges 6x. For a summarization workload that reads a lot and writes a little, a high output multiple barely matters; for a generation workload that writes long documents, it dominates. Model the ratio against your own input-to-output split rather than comparing input prices alone. Our LLM cost optimization guide covers how to profile that split.
Teams routinely find their actual spend running well above what the published rates implied. Four mechanics account for almost all of the gap, and none of them are hidden fees; they are properties of how these systems work.
This is why observability at the request level matters more than a monthly total. If you cannot see tokens per feature, per model, and per team, you cannot tell growth from waste, a problem we cover in our guide to AI cost visibility.
In order of impact, these are the changes that reliably cut an LLM bill. The first alone usually outperforms the other three combined.
Send each request to the smallest model that clears your quality bar, and escalate only what fails. Most production traffic is routine: classification, extraction, routing, short summaries, simple chat. Those run perfectly well on Haiku 4.5, Nova Lite, Flash-Lite, or GPT-5-nano at a fraction of flagship rates. Reserve Opus 5, GPT-5.5, and the frontier tier for the genuinely hard minority. Teams that introduce routing typically cut spend by 60 to 80 percent with no measurable quality loss, because they were paying frontier rates for work that never needed them.
Every major provider now offers prompt caching, which stores a repeated prefix, typically your system prompt, few-shot examples, or a fixed document, and charges a fraction of the input rate when it is reused. The standard discount is around 90 percent, and Claude Fable 5.1 goes further at 97.5 percent off input for cache reads. If your application sends the same 4,000-token system prompt on every call, caching it is the single easiest structural saving available, and it requires no change to model or quality.
Batch APIs process requests asynchronously, usually within 24 hours, at half price across OpenAI, Anthropic, Google, and Amazon Bedrock. Evaluations, backfills, enrichment, classification of historical data, and report generation almost never need a synchronous response. Moving that traffic to batch halves its cost for the price of a scheduling change.
Because output costs three to five times input and reasoning tokens bill as output, controlling generation length is disproportionately effective. Set maximum output tokens explicitly rather than leaving the default, instruct models to be concise where verbosity adds nothing, and tune the thinking or effort budget down on tasks that do not need deep reasoning. These are prompt-level changes that cost nothing to implement. Our guide to token budgeting and AI cost control covers how to set and enforce those limits in production.
Consider a support-triage feature handling 500,000 requests a month, each with about 2,000 input tokens and 300 output tokens. That is 1 billion input tokens and 150 million output tokens a month.
On Claude Fable 5.1 at $10 and $50, that costs $10,000 in input and $7,500 in output, about $17,500 a month. On Claude Sonnet 5 at $2 and $10, the same volume costs $2,000 and $1,500, about $3,500. On Claude Haiku 4.5 at $1 and $5, it costs $1,000 and $750, about $1,750. Triage is a classification task, so Haiku handles it well, and the frontier model would deliver no better outcome for ten times the price.
Now apply the other levers to the Haiku version. If 1,500 of those 2,000 input tokens are a fixed system prompt, caching them cuts input cost by roughly 65 percent overall. If half the volume is historical backfill that can run asynchronously, batch halves that portion. The $17,500 frontier bill and the roughly $900 optimized bill run the same feature at the same quality, and the difference is entirely decisions rather than technology.
Routing down is not always right, and the honest version of this advice includes the cases where paying more is cheaper overall. A frontier model that solves a task in one attempt can beat a cheap model that needs four attempts plus human correction, and the human time is usually the larger cost. Complex agentic coding, long-horizon multi-step work, high-stakes reasoning where errors are expensive, and tasks where a cheaper model demonstrably fails are all cases where the premium earns its place.
The test is empirical rather than ideological. Run the task on the cheaper tier, measure the failure rate and what failures cost you, and compare that to the price difference. Teams that skip this test tend to err in one direction or the other permanently: either paying frontier rates for classification, or forcing hard reasoning through a budget model and absorbing the error rate downstream.
Token pricing in 2026 is not really a pricing problem; it is a design problem with a price attached. The rate card spans 300 to one, the levers are documented and free to use, and almost every expensive AI bill traces back to the same two decisions: one model chosen for everything, and nobody measuring output tokens. Fix those and most of the cost takes care of itself.
So profile your input-to-output split, route by difficulty, cache the parts that repeat, batch the parts that can wait, and cap the parts that ramble. Then attribute what remains to the teams and features that generate it, because a bill nobody owns is a bill nobody optimizes. That attribution step is the subject of our guide to AI unit economics, and the wider practice lives in FinOps for AI.
LLM token costs range from about $0.035 per million input tokens on the cheapest models, such as Amazon Nova Micro, to $10 per million on frontier tiers like Claude Fable 5.1 and GPT-6 Astra. Output typically costs three to five times input. Mainstream production models cluster between $1 and $5 per million input tokens.
Amazon Nova Micro is the cheapest widely available model at about $0.035 input and $0.14 output per million tokens, followed by Nova Lite and GPT-5-nano. Among frontier-class models, DeepSeek V4 Flash is the lowest cost at about $0.22 and $0.66 off-peak. Cheapest per token is not always cheapest per task, since weaker models may need more attempts.
Generating text is more computationally demanding than reading it, so providers price output at roughly three to five times input across the board. This matters because reasoning or thinking tokens are also billed as output, meaning a short visible answer can carry a large hidden output cost on reasoning-heavy models.
Prompt caching stores a repeated input prefix, such as a system prompt or fixed document, so subsequent calls that reuse it are billed at a fraction of the standard input rate. The typical discount is around 90 percent, and Claude Fable 5.1 offers 97.5 percent off input for cache reads. It is usually net positive after a single reuse.
Route each request to the smallest model that meets your quality bar, which typically saves 60 to 80 percent on its own. Then cache repeated prompt prefixes for around 90 percent off that portion, move delay-tolerant work to batch APIs for 50 percent off, and cap maximum output tokens and thinking budgets, since output costs several times more than input.
Yes. OpenAI, Anthropic, Google, and Amazon Bedrock all offer batch processing at roughly 50 percent of standard rates, in exchange for asynchronous completion usually within 24 hours. It suits evaluations, backfills, enrichment, and classification of historical data, and requires no change to model or prompt.
Four mechanics usually explain the gap: reasoning tokens billed as output that the user never sees, conversation history resent as input on every turn, agentic workflows making dozens of model calls per user request, and failed calls and retries that still bill for the tokens they consumed.