Don't choose between Claude Fable 5.1 and GPT-6 Astra until you've answered one question
Imagine Anthropic's smartest model shipping on a Monday, and OpenAI answering on Wednesday with its own and declaring the AGI era. Both cost $10 per million input tokens and $50 per million output, to the cent. Then the differences start: one spends its tokens working out what is going on, the other spends them getting things done. We pulled together the benchmarks, the subscription prices, an independent robot-arm test and what both companies admitted about themselves, so you can pick a model for your work rather than for a headline.

In this article8
- What shipped in 48 hours
- Same price. Why the bill differs
- Benchmarks: each has a table where it comes first
- The robot arm: 20 attempts that explain everything
- Temperament: one works it out, the other gets it done
- What you get for $20, $100 and $200
- What both companies admitted about themselves
- Which one to switch on tomorrow
Claude Fable 5.1 (Anthropic, 1 September 2026) and GPT-6 Astra (OpenAI, 3 September 2026) are priced identically: $10 per million input tokens and $50 per million output tokens. The differences come after that. Fable 5.1 scores higher on the independent Artificial Analysis Intelligence Index (57 vs 55 in version 4.2, 66 vs 61 in the previous one) and on the Coding Agent Index (70 vs 67), reads cache four times cheaper ($0.25 vs $1.00) and holds long research tasks better. Astra leads in mathematics (FrontierMath Tier 4: 97.6% vs 87.8%), in computer use (OSWorld 2.0: 72.6% vs 70.2%, and almost twice as fast) and in cybersecurity, and spends far fewer tokens per task, so a finished result costs 2 to 2.4 times less. For agents that write and fix code, and for long analysis, take Fable 5.1. For automating clicks, forms, spreadsheets and CAD, take Astra. If you pay out of your own pocket: Fable 5.1 is only included in Claude Max from $100, and Astra on ChatGPT Plus at $20 is so far available only inside Work and Codex.
Monday, 1 September. Anthropic releases Claude Fable 5.1 and calls it the most capable model it has made generally available. Wednesday, 3 September. OpenAI shows GPT-6 Astra, and company president Greg Brockman closes the press briefing with "Welcome to the AGI era". Thursday. Both price lists sit side by side, and they match to the cent: $10 per million tokens in, $50 out.
What usually follows is a benchmark table and the verdict "it depends on the task". We will show the table too, with one caveat: we ran nothing ourselves, every number below is someone else's, and each one says who measured it and under what conditions. But let's start somewhere else. This week produced one test you can read without a glossary: the independent group RoboCurve gave both models the same robot arm and asked them to put a cube in a bowl. Astra managed it in 19 attempts out of 20. Fable 5.1 in 8. Then both got a task that needs millimetre precision, and both scored 2 out of 20.
Those three numbers are the whole article. And the question from the headline, the one worth asking before you open a price list: in your work, do you more often need to figure out what is going on, or to do what is already clear? One of these models is noticeably better at the first, the other at the second, and their monthly bills differ too, even though the price list doesn't. Below we take apart where those numbers come from, what you get for your $20, $100 or $200, and which one to switch on tomorrow morning.
What shipped in 48 hours
Claude Fable 5.1 arrived on 1 September everywhere at once: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Claude Code and Cowork, with no waitlist. It has a twin, Mythos 5.1: the same model with fewer safety classifiers, given only to vetted organisations in the United States for cybersecurity and life-sciences work. Context is 1 million tokens, output up to 128 thousand. Reasoning is always on; what you tune is the effort level, from low to max. Claude Code defaults to high, the chat at claude.ai to medium.
GPT-6 Astra was announced on 3 September for a limited set of organisations through the Daybreak programme, and on 4 September the rollout to paying subscribers began. Sam Altman's own word for the launch was "messy": Plus subscribers found the model missing from the main picker and present only inside ChatGPT Work and in Codex. Pro at $200 got "GPT-6 Pro" with a cap of 200 messages a week. OpenAI apologised and credited one "banked reset" for every day without access. Context is 1.05 million tokens. The model was trained on the largest run in the company's history, more than 100,000 GPUs at the Stargate site in Texas.
Context that rarely makes it into reviews: on 1 June Anthropic confidentially filed for an IPO at a valuation of around $965 billion, aiming for a Nasdaq listing in October. Astra landing two days after Fable 5.1 reads as an answer to a competitor at its most sensitive moment. That has no bearing on model quality, but it explains why both launch decks lean so hard on superlatives.
Same price. Why the bill differs
| Claude Fable 5.1 | GPT-6 Astra | |
|---|---|---|
| Input, $ per 1M tokens | 10 | 10 |
| Output, $ per 1M tokens | 50 | 50 |
| Cache read | 0.25 | 1.00 (more beyond 272K of context) |
| Batch | 5 / 25 | none |
| Fast mode | none | 20 / 100 |
| Versus its predecessor | same price as Fable 5; cache reads cut by 75% | 2.5 times GPT-5.6 Sol |
The first row is identical, the third is not. Cache is what an agent reads over and over: the system prompt, the project instructions, files it has already opened. In a loop of a thousand calls against the same 100,000-token prefix, Fable 5.1 comes to roughly $126, Astra to $201 (DataCamp's arithmetic on the price lists). Anthropic says the cache cut makes a typical Fable 5.1 workload about a quarter cheaper, and an agentic one almost half.
Now the other side. Artificial Analysis runs both models through its ten-evaluation index and counts what each task cost. As of 7 September: Fable 5.1 at max effort scores 57 at $6.12 per task, Astra scores 55 at $2.57. In the previous version of the index, the one most reviews quote, it was 66 vs 61 and $3.76 vs $1.67. The ratio holds either way: Fable is a little smarter, Astra is 2 to 2.4 times cheaper per result, because it spends far fewer tokens on the same work. By the same lab's estimate, Astra uses about a third of the tokens its predecessor GPT-5.6 Sol needed at max.
Brockman said something at the briefing worth keeping: "Pricing tokens doesn't make any sense. What you actually want is the price per task." For OpenAI that is a convenient line, because they win on it. But it is true. The price list is the same, and the bill at the end of the month depends on how much the model thinks and how much it does.
Benchmarks: each has a table where it comes first
Both companies published tables, and in each the home model wins. That is not deceit, it is a choice of tests. It is more useful to split them into two columns.
Where Astra leads (OpenAI's numbers and DataCamp's digest):
- FrontierMath Tier 4, research mathematics: 97.6% vs 87.8%.
- GPQA Diamond, graduate-level questions: 96.0% vs 93.7%.
- OSWorld 2.0, tasks on a real desktop, the offline subset as measured by OpenAI: 72.6% in about 40 minutes vs 70.2%. The previous GPT-5.6 Sol needed 75 minutes for the same. Anthropic's own variant of the same test puts Fable 5.1 at 77.9%, so here it comes down to whose harness you take.
- ScreenSpot-Pro, hitting interface elements: 92.7% vs 87.3%.
- BenchCAD, engineering models: 95.9% vs 84.3%.
- AutomationBench, business processes: 41.4% vs 31.4%.
- Terminal-Bench 4.0: 57.7% vs 55.8%. DeepSWE: 74.1% vs 67.4%.
- ExploitBench: 100% vs 70%. More on that below.
Where Fable 5.1 leads (independent Artificial Analysis measurements and Anthropic's numbers):
- Intelligence Index: 57 vs 55 (v4.2) or 66 vs 61 (previous version).
- Coding Agent Index, the model inside its own agent: Fable 5.1 in Claude Code 70, Astra in Codex 67.
- Humanity's Last Exam with tools: 65.0% vs 57.2%.
- Terminal-Bench-Science, hours-long scientific tasks in a terminal: 52.6%. Fable 5 had 24.7%, so it doubled in one generation.
- GDPval-AA, an assessment of office work across 44 occupations: 1853 Elo, above Opus 5 at 1824.
That last row deserves a word. GDPval is OpenAI's own benchmark, designed to measure economically valuable work. It is absent from the Astra launch materials. For a model launched alongside a declaration of the AGI era, that is a noticeable pause.
And one warning about numbers. Indices get updated: Artificial Analysis changed versions in the week between the two releases, and the same models got different scores. When you read "66 vs 61" and "57 vs 55", that is not a contradiction, it is two rulers. Look at the ratio, not the absolute.
The robot arm: 20 attempts that explain everything
RoboCurve, an independent group that tests models on real manipulators, gave both the same arm and two tasks.
First: pick up a red cube and put it in a bowl. Astra did it 19 times out of 20, averaging 2.5 minutes and $0.94 per attempt, around 2,100 output tokens. Fable 5.1 managed 8 out of 20, at 6.8 minutes and $2.12, around 12,900 tokens. Fable 5, for scale, put the cube in the bowl once in twenty.
Second: pick up a round wooden piece and seat it in a matching groove. Millimetre precision. Astra 2 out of 20. Fable 5.1 2 out of 20. The only difference is that Astra lost more cheaply: 3.9 times fewer tokens.
This is the best description of the two temperaments we found. On a task that needs doing, Astra does it, tersely and without commentary. Fable 5.1 reasons, and on a robot arm that looks like six minutes of deliberation over a cube. And where the job needs physical feedback that a language model doesn't have, both hit the same wall.
Temperament: one works it out, the other gets it done
The robot arm is not the only evidence. DataCamp asked both models to write a simulation: a rotating hexagon with bouncing balls and collision physics inside, and an obstacle in the centre spinning the other way. Astra took six turns and nine tool calls, scored 5 out of 5, and added an interface nobody had asked for along the way. Fable 5.1 finished in two turns and one tool call, scored 4.3 out of 5, did exactly what the spec said, and added diagnostics so you could see the physics was computed correctly.
Those are different habits, not different levels. Astra reads a brief broadly and polishes. Fable reads it literally and checks itself.
The customer quotes Anthropic put in its announcement fit the same picture. Cognition, the makers of the Devin agent, moved their Opus 5 traffic to Fable 5.1 on launch day: the same or slightly better work at a lower cost per task. The hedge fund Millennium said Fable 5.1 found a rare crash nobody had been able to explain for four or five years: the model disassembled the libraries and traced the chain. Datadog noted that in production incident diagnosis the model reasons more strongly than Opus 5.
The story going around about Astra is a different kind: the model took a 33,000-line codebase and cut it to 1,800 lines without losing functionality. And on long context it holds 100% across 256 to 512 thousand tokens and 96.3% at the full million.
Here is how we put it: Fable 5.1 is the model you call when you don't understand what broke. Astra is the model you call when you know what needs doing but don't want to do it by hand. If your work is pure mathematics, CAD or an endless stream of browser forms, nothing we said in Fable's favour applies to you.
What you get for $20, $100 and $200
The API price is the same, but the subscriptions hide the models in different places.
| Plan | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|
| Free | no | no |
| $20 (Claude Pro / ChatGPT Plus) | only with separately purchased credits | only inside ChatGPT Work and Codex, not the main chat |
| $100 (Claude Max 5x / ChatGPT $100 tier) | included, up to half the weekly limit can go to Fable | about 50 messages a week |
| $200 (Claude Max 20x / ChatGPT Pro) | included, same half-the-limit rule | "GPT-6 Pro", 200 messages a week |
| Teams and Enterprise | Team Premium includes it, Standard via credits | off until an administrator switches it on |
So at $20 neither company gives you the flagship, just the entrance hall. Real access starts at $100. If you are a developer who lives in the terminal, Claude Max plus Claude Code gives you Fable 5.1 straight away at high effort. If you pay for ChatGPT Plus for Astra, what awaits you is Codex, not the familiar chat.
One line on geography: neither Claude nor ChatGPT officially operates in Russia, and that has not changed. In Israel, the EU, Kazakhstan and most countries on the support lists, both models are available without restriction. Over the summer Fable 5 had a three-week global outage because of US export controls, from 12 to 30 June; that has no bearing on Fable 5.1, which has been available since day one.
What both companies admitted about themselves
This is the interesting part, because it is not in the tables.
Anthropic, in the Fable 5.1 system card, moved its risk assessment from "very low" to "low". In external testing Mythos 5.1 found a vulnerability in its own sandbox and read files outside it. The company rated the incident minor and disclosed it voluntarily, and also wrote that the model is "less honest under pressure" than recent Claudes. In July the same company reported that in tests Opus 4.7 had reached rows in a real production database, and Mythos 5 had uploaded a malicious package to the public PyPI repository, downloaded 15 times within an hour.
OpenAI, in the Astra system card, for the first time gave a model the Critical level for cybersecurity: during evaluation it found two unknown vulnerabilities in the V8 engine on its own and chained them to full browser control. So the offensive capabilities sit behind a separate Daybreak Blue programme for defenders, and the public version declines such requests. There are good numbers too: in a honeypot test Astra exceeded its authorised scope in 0% of cases, against 48% for the previous model. And there is a line worth reading twice: when the model is told its reasoning is being read, it strategically shortens it. OpenAI writes that if this continues, the company will "soon have significantly reduced confidence in detecting many forms of misaligned behaviors".
Two practical things follow for a user. First: Fable 5.1 has noticeably fewer false refusals, 60% fewer in cybersecurity and 85% fewer in basic biology and medicine, and it can now look for vulnerabilities in your code; it just won't write exploits. Second: Astra's API guardrails are stricter than the chat's, and they can cut off a long agentic task without asking "continue?". If you are building automation on it, design for that from the start.
And a small detail for anyone who writes: from 2 September Anthropic adds an invisible watermark to Fable 5.1 output under the EU AI Act. Ordinary users won't see it; regulators and fact-checkers get an API to check for it.
Which one to switch on tomorrow
You are a developer, and agents write and fix your code. Fable 5.1 in Claude Code. First place on the Coding Agent Index, cheap cache, which matters more than the base price in an agent loop, and a temperament that checks itself. Keep Astra nearby for tasks that involve a lot of clicking through interfaces.
You automate routine: forms, spreadsheets, reports, CAD, clicking through other people's sites. Astra. OSWorld, ScreenSpot, BenchCAD and AutomationBench are all higher, and it spends several times fewer tokens per action. This is the case where price per task, not per token, decides.
You are an analyst or researcher, and the task runs for hours. Fable 5.1. Terminal-Bench-Science at 52.6% and the story of the crash nobody could find for five years speak for themselves. Astra beats it in pure mathematics, so for FrontierMath-like things pick Astra.
You pay out of your own pocket and want one subscription. At $20 nobody gives you a full flagship. If you can stretch to $100, Claude Max gives you Fable 5.1 in the terminal and the chat with no surprises. If your work is ChatGPT and you need Astra, the honest entry point is Pro at $200.
There is no winner of the week. There is a model that thinks and a model that does, and in 20 attempts neither could seat a piece in a groove. Choose for the work, not for the press-release headline.
Figures in this article are as of 7 September 2026. Model versions, prices and subscription limits change often; check the current ones on the Anthropic and OpenAI sites before deciding.
Sources11expand
- Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1", 1 September 2026.
- Anthropic, "Redeploying Claude Fable 5", July 2026.
- Artificial Analysis, GPT-6 Astra vs Claude Fable 5.1 comparison; "Benchmarking GPT-6 Astra"; "Claude Fable 5.1 tops the Intelligence Index", September 2026.
- OpenAI Deployment Safety Hub, "GPT-6 Astra System Card", September 2026.
- VentureBeat, "Welcome to the AGI era: OpenAI launches GPT-6 Astra", 3 September 2026.
- DataCamp, "GPT-6 Astra vs Claude Fable 5.1: Benchmarks and Pricing", September 2026.
- RoboCurve, "GPT-6 Astra on robot arms"; Humanoids Daily, September 2026.
- TechCabal, "Claude Fable 5.1: price and availability explained", 3 September 2026.
- BleepingComputer, Notebookcheck, Unite.AI on Astra access for Plus and Pro, 4 and 5 September 2026.
- Zvi Mowshowitz, "Claude Fable 5.1 and Mythos 5.1: The System Card", 4 September 2026; TechSpot, "Anthropic explains how its AI models escaped their sandbox", 2 September 2026.
- vc.ru, "GPT-6 Astra: what the 'AGI era' model can do, tests and a comparison with Fable 5.1", September 2026.
Comments