Loading...


Updated 21 Sep 2026 • 7 mins read

Claude Fable 5.1 is Anthropic's highest self-serve tier at $10 per million input tokens and $50 output, unchanged from Fable 5, but with cache reads cut to $0.25 per million. This guide covers the full rate card, how the caching change alters agentic economics, comparisons, and when the premium pays off.
When Anthropic released Claude Fable 5.1 on September 1, 2026, the headline rates did not move at all. Input stayed at $10 per million tokens, output at $50, exactly what Fable 5 charged. On a pricing page, that looks like a non-event. In practice it was one of the more consequential pricing changes of the year, because the number that did change was the one most people ignore: the cache read rate fell by 75 percent, from $1.00 to $0.25 per million tokens.
That single line moves the economics of an entire class of workload. Every other Claude model charges 10 percent of the input rate for a cache hit; Fable 5.1 charges 2.5 percent. For long-running agents that re-read the same context on every step, the cheapest part of the bill just got four times cheaper, which is exactly the shape of workload the frontier tier exists for. This guide covers what Fable 5.1 actually costs, how the caching change alters the math, how it compares to the tier below, and the honest test for whether you should be using it at all.
The short answer Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, the same rates as the Fable 5 it replaces. Its distinguishing feature is cache pricing: cache reads cost $0.25 per million tokens, just 2.5 percent of the input rate, where every other Claude model charges 10 percent. Anthropic estimates this cuts typical Fable workloads by around 25 percent and cache-heavy agentic workloads by up to around 45 percent. The Batch API halves both input and output. Fable 5.1 is Anthropic's highest self-serve tier, priced at exactly twice Claude Opus 5, and it is worth the premium only where Opus 5 demonstrably falls short.
Fable 5.1 only makes sense in the context of the tier below it, so here is the complete current lineup. All figures are standard first-party API rates per million tokens in US dollars, as of September 2026.
| Model | Input / 1M | Output / 1M | Cache read / 1M | Positioning |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1.00 | $5.00 | $0.10 | Fast, low-cost: routing, extraction, classification |
| Claude Sonnet 5 | $2.00 | $10.00 | $0.20 | Default for most production workloads |
| Claude Opus 5 | $5.00 | $25.00 | $0.50 | Complex agentic coding and enterprise reasoning |
| Claude Fable 5.1 | $10.00 | $50.00 | $0.25 | Highest self-serve tier; long-horizon agentic work |
Read that cache column carefully, because it contains the anomaly. Cache reads normally scale with the model: the more expensive the model, the more expensive the cache hit. Fable 5.1 breaks the pattern, charging less for a cache read than Opus 5 does despite costing twice as much per input token. That is not an error in the table; it is the deliberate design of the release.
One other current-pricing note worth recording, because it reversed an earlier announcement: Claude Sonnet 5 launched at $2 and $10 as introductory pricing due to end on August 31, 2026, and Anthropic has since confirmed those are now the standard rates. The scheduled increase to $3 and $15 will not happen. Our Claude pricing guide covers the full lineup and its history.
Prompt caching stores a repeated prefix of your input, typically a system prompt, a set of examples, a codebase, or a long document, so that subsequent calls reusing it are billed at the cache read rate instead of the full input rate. Writing to the cache costs slightly more than standard input, usually 1.25 times, so caching becomes net positive after a single reuse and dramatically positive after many.
The size of that benefit depends entirely on the cache read discount, and this is where Fable 5.1 separates from everything else. On a standard Claude model, a cached token costs 10 percent of input. On Fable 5.1 it costs 2.5 percent. Put concretely: reusing a 100,000-token context costs $1.00 on Fable 5, $0.50 on Opus 5, and $0.25 on Fable 5.1. The most expensive model in the lineup is now the cheapest place to re-read a large fixed context.
That is why Anthropic frames the saving in workload terms rather than rate terms. Typical Fable workloads see around 25 percent lower cost, and highly agentic, cache-heavy workloads up to around 45 percent, because those workloads are dominated by repeatedly re-reading the same context across many steps. An agent that reads a large codebase on every one of forty planning steps is paying the cache read rate forty times, and quartering that rate changes the total far more than any change to the headline input price would.
Rate cards are abstract, so consider a realistic long-horizon agent task: a coding agent working across a large repository, which loads 150,000 tokens of context, runs 30 reasoning and tool steps, and generates 40,000 output tokens in total.
Without caching, that context is re-sent on every step: 150,000 times 30 is 4.5 million input tokens, costing $45 at Fable 5.1 rates, plus $2 for output, roughly $47 per task. With caching, the context is written once (about $1.88 at the 1.25x write rate) and read 29 times at $0.25 per million, roughly $1.09, plus the same $2 of output. The task lands near $5. The same task on Fable 5, with its $1.00 cache reads, would have cost about $8.25, which is where the 40-percent-class saving comes from.
Two lessons fall out of that arithmetic. First, on any context-heavy workload, caching is not an optimization to consider later; it is the difference between a viable feature and an unaffordable one. Second, comparing frontier models on input price alone is misleading, because for this shape of work the cache rate drives the bill and the input rate barely appears.
The most important comparison is internal. Fable 5.1 costs exactly twice Opus 5 on both input and output, and Anthropic itself positions Opus 5 as the model most teams should start with. That is an unusually candid signal from a vendor, and it should be taken at face value: the frontier tier is not a general upgrade, it is a specialist tool. A task that Opus 5 completes correctly gains nothing from Fable 5.1 except a doubled bill.
Where the premium earns its place is long-horizon agentic work, tasks whose difficulty causes cheaper models to fail or loop, and context-heavy workloads that benefit from the cache rate. In that last category the comparison can even invert: a heavily cached agentic workload can cost less on Fable 5.1 than on Opus 5, because the cache reads that dominate the bill are half the price, even though the input rate is double.
Externally, Fable 5.1 sits at exactly the same headline rates as OpenAI's GPT-6 Astra, $10 and $50, which makes the two the joint-most-expensive self-serve models available. The differentiators are elsewhere: Astra is positioned around computer use and operating software, Fable 5.1 around long-horizon reasoning and cache-heavy agentic work. Our comparison of Anthropic and OpenAI covers the wider trade-off, and our guide to GPT-6 Astra covers that side in detail.
One point of frequent confusion is worth settling, because it affects procurement rather than price. Claude Mythos 5.1 and Claude Fable 5.1 are the same underlying model, priced identically. The difference is safeguards and access: Fable 5.1 carries additional safety measures in areas such as biology, cybersecurity, and model research, and is the version available through normal self-serve channels, while Mythos 5.1 has more permissive safeguards and is available only to a small number of vetted organizations through Anthropic's trusted access programs.
For nearly every commercial team the practical answer is that Fable 5.1 is the model you can buy, and Mythos is not a cheaper or more capable alternative to negotiate for. It is the same capability under a different access regime.
If you have concluded the tier is right for your workload, four levers reduce what it costs, in rough order of impact.
These are the same levers that govern any token bill, covered in our LLM cost optimization guide and guide to LLM token costs by model, but the unusual cache rate shifts their priority order on this particular model.
Claude Fable 5.1 is a frontier model priced like one, at $10 and $50 per million tokens, twice what Opus 5 charges at both ends. The interesting part of its pricing is not the headline but the footnote: cache reads at $0.25 per million, a quarter of what Fable 5 charged and half what Opus 5 charges, which makes the most expensive model in the lineup the cheapest place to repeatedly re-read a large context.
So the decision framework is narrower than the price tag suggests. If your work is context-heavy and agentic, running many steps over the same large body of information, Fable 5.1's economics are better than its rate card implies and occasionally better than the tier below. If your work is ordinary production traffic that Opus 5 or Sonnet 5 handles correctly, the premium buys nothing at all. Test on the cheaper tier first, measure where it genuinely fails, and reserve the frontier for those cases. That discipline is worth more than any rate card.
Claude Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, the same as the Fable 5 it replaced. Cache reads cost $0.25 per million, and the Batch API halves both input and output to $5 and $25. Confirm current rates on Anthropic's official pricing page.
The headline rates did not change; both cost $10 input and $50 output per million tokens. The change is cache pricing: Fable 5.1 cache reads cost $0.25 per million, down 75 percent from $1.00 on Fable 5. Anthropic estimates this reduces typical Fable workloads by around 25 percent and cache-heavy agentic workloads by up to around 45 percent.
Because Fable 5.1 prices cache reads at 2.5 percent of its input rate, while every other Claude model uses the standard 10 percent multiplier. That gives Fable 5.1 a $0.25 cache read against Opus 5's $0.50, even though Fable's input rate is twice as high. It is a deliberate design choice aimed at long-running, context-heavy agentic workloads.
Only where Opus 5 genuinely falls short. Fable 5.1 costs exactly twice Opus 5 at both ends, and Anthropic positions Opus 5 as the model most teams should start with. The premium earns its place on long-horizon agentic work, tasks where cheaper models fail or loop, and heavily cached context-heavy workloads, where the cheaper cache rate can offset the higher input price.
They are the same underlying model at identical prices. Fable 5.1 includes additional safeguards in areas such as biology, cybersecurity, and model research and is available through normal self-serve channels. Mythos 5.1 has more permissive safeguards and is available only to vetted organizations through Anthropic's trusted access programs.
Cache aggressively, since at $0.25 per million cache reads it is by far the highest-return lever on this model. Then use the Batch API for delay-tolerant work at half price, cap maximum output tokens and reduce thinking budgets because output costs five times input, and route routine steps to Haiku 4.5 or Sonnet 5, escalating to Fable only where needed.
Claude Sonnet 5 costs $2 per million input tokens and $10 per million output. It launched at those rates as introductory pricing through August 31, 2026, and Anthropic has since confirmed them as the standard price; the previously scheduled increase to $3 and $15 will not take effect.