AI/Comparison

Claude Fable 5.1 vs GPT-6 Astra: Price and Benchmarks Compared

Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra both launched this September — here’s how their price, benchmarks and coding chops stack up.

Claude Fable 5.1 and GPT-6 Astra are now the flagship models from Anthropic and OpenAI, and on paper they are closer than any pairing the two labs have shipped before: both charge $10 per million input tokens and $50 per million output tokens on their standard API tiers, both top out at 128,000 output tokens, and both claim to be the best model yet for serious coding and knowledge work. The real differences show up in the details — Astra pulls ahead on math, abstract reasoning and raw cybersecurity capability, while Fable 5.1 is markedly cheaper to run at scale because of a steep cut to its cached-token pricing, and it beats Astra on several agentic-reasoning benchmarks that lean on general judgment rather than pure problem-solving. Which one is "better" depends heavily on what you are optimizing for.

The short version

  • Release dates: Claude Fable 5.1 (and its lightly-safeguarded sibling Claude Mythos 5.1) shipped September 1, 2026. GPT-6 Astra began rolling out September 3 and reached general availability on ChatGPT Plus, Pro, Business and Enterprise on September 4, 2026.
  • Price: Identical headline rate — $10/M input, $50/M output — but Fable 5.1 cuts cached-input reads to $0.25/M (down from $1/M), while Astra's cached input is $1/M, rising to $2–$20/M on prompts over 272K tokens.
  • Context window: Fable 5.1 supports 1,000,000 tokens; Astra supports 1,050,000 tokens. Both cap output at 128,000 tokens.
  • Benchmarks: Astra leads on FrontierMath, GPQA Diamond, ARC-AGI-2/3, AutomationBench and cybersecurity evals. Fable 5.1 leads on Humanity's Last Exam (with tools) and the Artificial Analysis Intelligence Index, and edges Astra on some agentic terminal-coding scores.
  • Safety posture: Astra is the first OpenAI model to hit the "Critical" cybersecurity threshold under OpenAI's Preparedness Framework and refuses advanced exploit-creation tasks by default. Fable 5.1 can now find (but not weaponize) software vulnerabilities.

What Anthropic and OpenAI just shipped

Anthropic introduced Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, describing them as "the world's most advanced models for coding and knowledge work." The two are, in Anthropic's own words, the same underlying model with different levels of safeguards: Fable 5.1 is generally available to anyone with a paid Claude plan or API account, while Mythos 5.1 is restricted to vetted organizations through Anthropic's Cyber Verification and Life Sciences Verification programs, where lighter safeguards let it do things like penetration testing and exploit generation that Fable 5.1 will not.

OpenAI's GPT-6 Astra followed three days later, rolling out first to a limited set of organizations on September 3 before reaching all ChatGPT Plus, Pro, Business and Enterprise users, plus the OpenAI API, Microsoft Azure and AWS Bedrock, by September 4. OpenAI calls it "state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work," and it arrives as a bigger jump over GPT-5.6 Sol than Fable 5.1 is over Fable 5 — Astra saturates several benchmarks that previous models were nowhere close to, including a 99.9% score on ARC-AGI-3 and a 100% score on ExploitBench.

Both launches are pitched squarely at the same audience: developers and knowledge workers who need a model that can carry a long, tool-heavy task from start to finish without losing the thread. If you're weighing this pair against the generation before it, our earlier look at Gemini 3.8 Flash, GPT-6 and Claude Opus 5.5 is a useful baseline for how fast pricing and scores have moved in just a few weeks.

Pricing: Fable 5.1 vs GPT-6 Astra

Sticker pricing is identical. Both models charge $10 per million input tokens and $50 per million output tokens through their standard API tier. The real gap is in prompt caching, which is where most of the cost sits in agentic, tool-heavy workloads that repeatedly re-read the same context.

Anthropic cut Fable 5.1's cached-input price by 75%, to $0.25 per million tokens (2.5% of the base input rate, versus 10% on most earlier Claude models), specifically to bring down the cost of long agentic sessions. Anthropic says this reduces typical workload costs by about 25% relative to Fable 5, and by up to roughly 45% for complex coding and highly agentic tasks. GPT-6 Astra's cached input, by contrast, is priced at $1 per million tokens for prompts under 272K tokens — 10% of its base input rate — and both its input and cache rates double for anything longer, with output also rising to $75/M past that threshold, per OpenAI's published pricing. OpenAI also offers a Fast mode at 2x the standard price for lower latency, and Batch/Flex processing at half price for workloads that can tolerate delay.

Net effect: for short, one-shot prompts the two models cost the same. For the kind of long, cache-heavy agentic coding session that a tool like Claude Code or Codex runs all day, Fable 5.1 is meaningfully cheaper per completed task, because a much larger share of its token bill is cache reads billed at a quarter of the input rate rather than a tenth of it.

Spec / PriceClaude Fable 5.1GPT-6 Astra
ReleasedSeptember 1, 2026September 3–4, 2026
Input price (per million tokens)$10.00$10.00 (≤272K tokens); $20.00 above
Output price (per million tokens)$50.00$50.00 (≤272K tokens); $75.00 above
Cached input price (per million tokens)$0.25$1.00 (≤272K tokens); $2.00 above
Context window1,000,000 tokens1,050,000 tokens
Max output128,000 tokens128,000 tokens
Reliable knowledge cutoffJune 2026April 30, 2026
API model IDclaude-fable-5-1gpt-6-astra
Available viaClaude API, AWS Bedrock, Google Cloud, Microsoft FoundryOpenAI API, ChatGPT, Microsoft Azure, AWS Bedrock

Coding and agentic terminal benchmarks

Coding is the category both companies lead with, and it's the closest race of the whole comparison. On OpenAI's own published comparison table, GPT-6 Astra scores 57.9% on Terminal-Bench 4.0 against Fable 5.1's 55.8% — a real but modest lead. On the same table, Astra's FrontierCode 1.1 scores (64.5% Extended, 53.3% Main) also edge out Fable 5.1 (63.6% and 50.9%), and Astra's AutomationBench score of 41.4% is well clear of Fable 5.1's 31.4%.

Anthropic's own launch benchmarks, run before Astra existed, show Fable 5.1 improving sharply over Fable 5: Terminal-Bench 4.0 climbs from 42.0% to 55.8% (and to 60.9% for the less-restricted Mythos 5.1), and CursorBench 3.2.0 rises to 73.4% — a score Anthropic quotes engineering teams at Datadog and SpaceXAI praising specifically for the model's ability to verify its own work on multi-step coding tasks. Anthropic's Terminal-Bench-Science 0.1 score of 52.6% (versus 24.7% for Fable 5) speaks to a similar jump in agentic research and debugging ability, not just one-shot code generation.

Read together, the two companies' self-reported numbers point to Astra having a small, consistent edge on raw coding-benchmark ceilings, while Fable 5.1 has closed most of the gap that existed between Claude and GPT models a generation ago — and does so at a lower effective cost per long agentic session, thanks to its cache pricing. Anyone choosing between coding assistants built on either model should also weigh the surrounding tooling; our comparison of Cursor, GitHub Copilot, Claude Code and Codex covers how those harnesses differ on top of the underlying model.

Knowledge work, reasoning and math

This is where the results genuinely split. GPT-6 Astra's biggest published gains are in the categories OpenAI leans on for its "AGI-era" framing: it saturates FrontierMath Tier 4 at 97.6% (versus 87.8% for Fable 5.1), scores 96.0% on GPQA Diamond (Fable 5.1: 93.7%), and posts 99.9% on ARC-AGI-3 and 95.0% on ARC-AGI-2, both categories where Anthropic hasn't published a comparable Fable 5.1 number.

Fable 5.1 comes out ahead on the harder, tools-enabled version of Humanity's Last Exam — 65.0% versus Astra's 57.2%, according to OpenAI's own comparison table — and on the Artificial Analysis Intelligence Index v4.1.1, an aggregate score that OpenAI's table lists Fable 5.1 at 65.7 against Astra's 61.2. Anthropic's own figures also show Fable 5.1 well ahead of GPT-5.6 Sol on GDPval-AA v2, a knowledge-work benchmark, at 1,853 versus 1,711. In practice, that split suggests Astra is the stronger pure problem-solver on closed-form math and science questions, while Fable 5.1 is more consistent on open-ended reasoning tasks that mix judgment with tool use — the kind of work that dominates day-to-day knowledge work rather than benchmark leaderboards.

BenchmarkClaude Fable 5.1GPT-6 AstraGemini 3.8 Flash
Terminal-Bench 4.0 (agentic coding)55.8%57.9%19.1%
FrontierCode 1.1 Main50.9%53.3%43.6%
FrontierMath Tier 487.8%97.6%—
GPQA Diamond93.7%96.0%95.3%
Humanity's Last Exam (with tools)65.0%57.2%—
ARC-AGI-290.0%95.0%—
AutomationBench31.4%41.4%—
HealthBench Professional (length-adjusted)58.1%63.4%52.1%
Artificial Analysis Intelligence Index v4.1.165.761.258.7

Figures above come from OpenAI's own published GPT-6 Astra benchmark tables, which independently re-ran Claude and Gemini models under OpenAI's evaluation harness — a different setup from Anthropic's self-reported numbers elsewhere in this article, so treat cross-vendor comparisons as directional rather than exact.

Computer use and agentic speed

OpenAI put unusual emphasis on computer-use performance for Astra, and the numbers back that framing up: on OSWorld 2.0's offline benchmark, Astra scores 72.6% in roughly 40 minutes per task, versus GPT-5.6 Sol's 65.7% in about 75 minutes — a real accuracy gain delivered in about 47% less time. OpenAI also reports Astra completing tasks 1.9x faster than GPT-5.6 Sol on the Mind2Web benchmark when paired with an updated Codex harness, and a ScreenSpot-Pro score of 92.7% without tool assistance.

Anthropic doesn't publish a directly comparable Astra-era OSWorld number for Fable 5.1, but its own tables show Fable 5.1 improving over Fable 5 on the same benchmark, from 72.9% to 77.9% (partial scoring) and from 36.1% to 41.7% (strict scoring). Both companies are converging on the same idea — that raw intelligence matters less for computer-use tasks than the ability to work fast and stay oriented across a long session without losing context, which is also why both launches spent as much time on agentic harness changes (Codex's context "notes" for Astra, safeguard tuning for Fable 5.1) as on the underlying model weights.

Cybersecurity, safety and what each model won't do

The most consequential difference between the two launches isn't a benchmark score — it's what each company decided to hold back. GPT-6 Astra is the first OpenAI model to meet the "Critical" threshold for cybersecurity under OpenAI's Preparedness Framework: it hit a perfect 100% on ExploitBench (versus 78.5% for GPT-5.6 Sol) and 88.0% on SRE-Bench reverse-engineering tasks when tested without production safeguards. In the shipped product, Astra refuses more advanced cybersecurity tasks such as creating proof-of-concept exploits, with OpenAI planning to loosen those restrictions gradually for vetted defenders through its Daybreak program.

Anthropic took a narrower step in the same direction: Fable 5.1 can now be used to discover software vulnerabilities — defensive security-scanning work — but not to develop exploits for them, and Anthropic reports its updated safeguards cut false-positive interventions in the cybersecurity domain by about 60%. The more capable, fewer-guardrails version of the same model, Mythos 5.1, is where Anthropic puts the exploit-generation and penetration-testing capability, gated behind its Cyber Verification Program for vetted US organizations rather than shipped broadly. Functionally, both companies landed on a similar answer — the most dangerous capabilities exist in the model but are locked behind separate access programs — they just drew the line for the general-release model in slightly different places.

How to access each model today

Claude Fable 5.1 is generally available now: anyone with a paid Claude.ai plan or a Claude API account can use it, and it's also live on AWS Bedrock, Google Cloud Vertex AI and Microsoft Foundry as claude-fable-5-1, per Anthropic's models overview. Claude Mythos 5.1 is not publicly available — access requires applying through Anthropic's Cyber Verification Program (for cyberdefense) or Life Sciences Verification Program, and is currently limited to US organizations while Anthropic works with the government to widen eligibility.

GPT-6 Astra reached ChatGPT Plus, Pro, Business and Enterprise users by September 4, alongside the OpenAI API, Microsoft Azure and AWS Bedrock as gpt-6-astra, listed in full on OpenAI's model documentation. In ChatGPT it appears as GPT-6 Pro on Pro, Business and Enterprise tiers, and is included within existing subscription allowances, with the option to buy additional usage credits. Enterprise workspace admins must turn Astra on for their organization; it's off by default at launch.

Which one should you use

If your work is dominated by closed-form problems with a checkable answer — competition math, dense scientific literature, formal cybersecurity research, or anything that benefits from OpenAI's ARC-AGI and FrontierMath gains — GPT-6 Astra's benchmark ceiling is higher, and its computer-use speed improvements are a genuine practical upgrade for browser- and desktop-automation tasks. It's also the only one of the two you can run inside ChatGPT's Business and Enterprise document, spreadsheet and slide tooling out of the box.

If your work looks more like long agentic coding sessions, multi-step research with a lot of re-read context, or general knowledge-work tasks where judgment matters as much as raw problem-solving, Fable 5.1's combination of a 65.0% Humanity's Last Exam (with tools) score, a leading Artificial Analysis Intelligence Index result, and cache pricing that's a quarter of the standard rate makes it the more cost-effective choice for sustained use — particularly inside Claude Code or any workflow that repeatedly re-reads the same large context. Teams that don't need Fable 5.1's newest gains at all might still find Anthropic's cheaper mid-tier model does the job; see our look at Claude Opus 5.5's price and benchmarks for that option.

What to watch next

Both families are still expanding. Anthropic has said it's bringing "many of the improvements" in Fable 5.1 to the rest of the Claude lineup, and Mythos 5.1 access is expected to widen beyond US organizations as Anthropic coordinates with the government on export terms. OpenAI, meanwhile, has already followed Astra with the smaller GPT-6 Sol and Luna models on September 22 and is expected to keep loosening Astra's cybersecurity restrictions for vetted Daybreak defenders in the coming weeks. Given how close the two flagships are on price and how quickly both companies are iterating, the more durable question for most buyers isn't which model wins today's benchmarks — it's which vendor's roadmap and safeguard model fits the workload you're actually running.

Frequently asked questions

Is Claude Fable 5.1 or GPT-6 Astra better for coding?

GPT-6 Astra scores slightly higher on most published coding benchmarks, including 57.9% versus 55.8% on Terminal-Bench 4.0, but Fable 5.1 is close behind and considerably cheaper for long, cache-heavy agentic coding sessions thanks to its $0.25 per million token cached-input price.

How much does GPT-6 Astra cost compared to Claude Fable 5.1?

Both charge $10 per million input tokens and $50 per million output tokens on their standard tier. The difference is caching: Fable 5.1's cache reads cost $0.25 per million tokens versus $1 (rising to $2) for Astra, and Astra's input and output rates roughly double on prompts over 272,000 tokens.

What is Claude Mythos 5.1?

Mythos 5.1 is the same underlying model as Fable 5.1 but with fewer safety safeguards, enabling capabilities such as exploit generation and penetration testing. It isn't publicly available; access requires approval through Anthropic's Cyber Verification Program or Life Sciences Verification Program.

Why does GPT-6 Astra refuse some cybersecurity requests?

Astra is the first OpenAI model to meet the "Critical" cybersecurity threshold under OpenAI's Preparedness Framework, so OpenAI restricts advanced tasks like creating proof-of-concept exploits by default, with plans to expand access gradually to vetted defenders through its Daybreak program.

What's the context window on each model?

Claude Fable 5.1 supports a 1,000,000-token context window; GPT-6 Astra supports 1,050,000 tokens. Both cap output at 128,000 tokens per response.

When did each model launch?

Claude Fable 5.1 and Claude Mythos 5.1 launched on September 1, 2026. GPT-6 Astra began rolling out on September 3 and reached general availability across ChatGPT and the API by September 4, 2026.

Sources

More on AI Model Benchmarks →AIClaudeAnthropicOpenAIGPT-6AI benchmarks
Theo Park
Written byTheo Park

Theo Park runs the AI desk at Pandromeda. He follows model launches from the frontier labs and the open-weight community, tracks the assistants and developer tools built on them, and explains what each release changes on pricing, capability and safety. His reporting leans on primary sources: model cards, technical reports, API documentation and the companies' own announcements.

More from AI

See all