Claude Fable 5.1 2026: Mythos 5.1, cache price, and the evals
Anthropic split one weight into public Fable 5.1 and gated Mythos 5.1. Name the SKU before reading the 52.6% or the AA 66.
Anthropic shipped two Claude 5.1 SKUs on 1 September 2026: a public Fable, and Mythos behind trusted access. Same underlying model, not a new flagship family.
The query is Claude Fable 5.1 and Claude Mythos 5.1. What follows matches the anthropic.com post and Artificial Analysis’s 1 September note. Vendor benches are not independent replications. Prior safety-taste note: TASTE / Fable 5.
What the launch post prints
Fable 5.1 is pitched for coding, knowledge work, and long-running problem solving. Claude Code defaults to High effort; Cowork and claude.ai default to Medium. List price is unchanged at $10 / $50 per million tokens. The change is cache reads at $0.25. Anthropic’s estimate uses four weeks of August usage: about 25% cheaper on typical loads, up to about 45% when cache reads dominate. The API name is claude-fable-5-1.
How to read the three rows
- 1
Name the public SKU before the gated one
Developers get Fable on the API, claude.ai, Claude Code, and the three clouds. Mythos needs a verification program — not the same subscription key. Anthropic also writes that Claude Security is now powered by Mythos 5.1, scanning codebases and proposing patches for human review.
- 2
The vendor science bench is not AA’s Terminal-Bench
Official Terminal-Bench-Science 0.1: Fable 5.1 at 52.6%. The public leaderboard (other harness) lists Opus 5 at 30.0% and Fable 5 at 21.4%; Anthropic’s setup reproduces 29.0% and 24.7%, both within noise. AA’s number is Terminal-Bench v2.1 at 91.4% max. Do not mash 52.6% and 91.4% into one crown. Tie-band: Muse Spark 1.3.
- 3
A cheaper cache is not a cheaper task bill
AA: 66 at max with default fallback, ahead of Opus 5 at 63, Fable 5 at 62, and GPT-5.6 Sol at 61. Fallback served about 4% of output tokens via Opus 4.8 or Opus 5. Cost is about $3.76 per index task versus $3.14 for Fable 5 — roughly 20% more — because output tokens are about 1.7×. The cache cut saves about $1.40; without it the task would be about $5.16. xhigh scores 65 at about $2.72.
Science agents and knowledge work
Anthropic writes GDPval-AA v2 at 1853 versus 1824 for Opus 5. AA repeats 1,853 Elo and notes overlapping confidence intervals with Opus. AutomationBench is 31.4%; CursorBench 3.2.0 is 73.4%. Vendor HLE is 60.9% without tools and 65.0% with tools; AA’s own HLE is 59.1%. Name the suite.
Lab science cases stay vendor narrative
Mythos used open-source protein tools; two external orgs validated designs: nearly 50% hit rate across 12 targets, and about 10× the best Adaptyv competition affinities on three targets. Fable built a higher-resolution elevation map for about a third of Venus from Magellan radar. Mythos wrote GPU kernels for seven open biology models; Anthropic writes up to about 2.5×. These are not a public reproduce pack.
Sharper gates are not a public pentest API
Anthropic writes about 60% fewer cyber false-positive blocks. Fable may now find software vulnerabilities; exploit generation, pentesting, and binary scanning still route to Opus. Enterprise Frontier Safeguards keep monitoring data in the customer’s cloud, phased this fall; eligible customers can use Fable 5.1 with zero data retention until then. Models released after 2 August carry an EU-practice invisible watermark.
| Claim | Checkable source | How to read it |
|---|---|---|
| Science 52.6%; TB 4.0 55.8% / 60.9% | anthropic.com 2026-09-01 | Vendor-run; science SE about ±3.5–4.5 |
| AA max 66; ~$3.76 / task | Artificial Analysis 2026-09-01 | Third-party index; default fallback ~4% tokens |
| Cache read $0.25; list $10 / $50 | Anthropic price card | Reads cheaper, output still dear; task bill separate |
Anthropic: Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards.
# Claude 5.1 · Anthropic 2026-09-01 · AA 2026-09-01
sku: claude-fable-5-1 # generally available
sku: claude-mythos-5-1 # trusted access
price: $10 / $50 per 1M
cache_read: $0.25 per 1M # was $1.00
tb_science_0.1: 52.6% # vendor; SE ±3.5–4.5
tb_4.0: 55.8% / 60.9% Mythos # vendor
aa_index_max: 66 # default fallback ~4%
aa_task_cost_max: $3.76Boundaries
- Vendor tables ≠ a third-party index
- Science 52.6% and HLE 65.0% are lab-named. AA 66 and AA’s HLE 59.1% are another harness. Do not mash them into one crown.
- The Mythos SKU ≠ a key anyone can buy
- Trusted access is for vetted defenders and life scientists. The looser gates are the access condition, not an open checkpoint.
- This page is not a meeting brief
- Release and eval only. For a short shared board, open wbmeet and create or join a room.
Questions worth checking
Are Fable 5.1 and Mythos 5.1 the same weights?
Anthropic writes “the same model” with different safeguard levels. Public docs do not publish a weight hash or downloadable checkpoint. Treat them as two SKUs with two access policies.
Is 52.6% or 66 the “official” score?
Neither is a regulator’s finding. 52.6% is Anthropic’s Terminal-Bench-Science. 66 is AA at max with default fallback. Name the suite and the effort row on any lab note.
Can a Pro or Team plan pick Mythos?
No. Mythos goes through verification programs and is currently limited to US organizations. CVP Mythos-class access is written as “near future,” with no full-open date.
Does this prove Anthropic passed other labs?
Not as a crown by itself. AA calls 66 the highest it has measured, but GDPval intervals overlap Opus, and the science bench carries a standard error. This page only unpacks the release.