Hands On: Reading Astra's launch benches
Alex Finn says GPT-6 Astra destroys every benchmark. Artificial Analysis has it tied with Sol and behind Fable 5.1. Here is how to read a launch video without confusing OpenAI's scoreboard with a public one.

Lab benches are not a neutral referee
When OpenAI (or any frontier lab) publishes or "leaks" a launch comparison, three things are usually true at once:
- The suite is curated. You see the benches where the new model is strong. You do not always see the ones where it is average, expensive, or behind.
- The protocol is theirs. Effort settings, scaffolds, tool access, retries, and which variant of a bench (ARC-AGI-1 vs ARC-AGI-2, SWE flavors, Terminal-Bench versions) can move a bar more than the model name on the axis.
- The audience is buyers and markets. Launch benches are closer to a product demo than to an audit. That does not make them fake. It makes them advocacy.
Finn is doing what a good YouTube launch video does: amplify the lab's best frame. OpenAI's frame this week is capability + Critical cyber + price-per-task. The hand-picked coding and ARC bars serve that story. They are not required to line up with Artificial Analysis, LMSYS, or anyone else's composite.
So if Finn says Astra is "much higher than everything," ask: higher on which scoreboard, assembled by whom?
What the public board we use actually says
Artificial Analysis runs GPT-6 Astra independently and publishes an Intelligence Index. As of 3 Sep 2026 on theaigentic.com/models:
| Model | AA Index | Effort setting scored | List $/1M (in / out) |
|---|---|---|---|
| Claude Fable 5.1 | 66 | max (adaptive) | $10 / $50 |
| Claude Mythos 5.1¹ | 66 | max (adaptive) | $10 / $50 |
| Claude Opus 5 | 63 | max | $5 / $25 |
| GPT-6 Astra | 61 | max | $10 / $50 |
| GPT-5.6 Sol | 61 | max | $4 / $20² |
| Grok 4.6 | 61 | high³ | $2 / $6 |
¹ Mythos 5.1 shares Fable 5.1's weights and score; it is available only to approved organizations, so for most readers Fable is the row that matters. ² Promotional pricing at time of writing. ³ Highest effort setting xAI exposes; AA has no "max" page for Grok 4.6. Scores across the table are the highest effort each vendor offers, which is the closest we can get to matched settings. ⁴ No Gemini row: on our Models cut, Google's best published Index score is Gemini 3.7 Flash at 56 (high), below this ≥61 band. Gemini 3.1 Pro Preview sits at 48.
Astra is among the leaders. It is not alone at the top, and it is not ahead of Fable 5.1 on that index. You can believe both statements without contradiction:
- On OpenAI's launch suite, Astra looks like a step-change on the benches they chose to show.
- On AA's public index, Astra is a strong frontier model in Sol's quality band, not a unanimous king.
How AA grading is done for a model like Astra
When we put 61 next to GPT-6 Astra, that number is not a vibe score and it is not OpenAI's deck. Here is the pipeline Artificial Analysis uses for a frontier API model, and what we copy onto the site.

1. Pick the variant they will score
OpenAI ships one model id (gpt-6-astra) with reasoning effort knobs (low → max). AA does not collapse those into one mystery number. It publishes separate pages — for example Astra (max), xhigh, and low. Our Models table cites max, the same way we cite Sol and Fable at their top published settings, so the column stays as comparable as the vendors allow.
2. Run the Intelligence Index suite independently
Index v4.1.1 is a nine-eval composite, not a single coding quiz:
| Eval in the Index | What it stresses |
|---|---|
| GDPval-AA v2 | Agentic real-world work |
| τ³-Banking | Agentic tool use |
| Terminal-Bench v2.1 | Agentic coding & terminal use |
| SciCode | Scientific coding |
| Humanity's Last Exam | Hard reasoning & knowledge |
| GPQA Diamond | Graduate-level science QA |
| CritPt | Physics reasoning |
| AA-Omniscience | Knowledge + hallucination tradeoff |
| AA-LCR | Long-context reasoning |
AA runs those itself (or under its published protocol). That is the main difference from a launch video: the lab does not get to choose only the benches where the new model spikes, and the harness is documented for the Index as a whole.
3. Weight the suite into one Index score
Each eval contributes to a single Intelligence Index number. Per AA's Index methodology, the score is a weighted average scaled 0–100, with four categories each contributing 25%: agents, coding, general capability, and scientific reasoning. For Astra max, AA currently reports 61. That places it roughly #8 among proprietary models on AA's full page, which lists more variants and vendors than our table above — our table shows only the top scores per model family. Low effort lands at 57 on AA's low page. Same model, same Index weighting, different effort — which is why we always show the setting next to the score.
4. Attach price, context, and verbosity — separately
AA also records list $/1M input and output, context window, and how many tokens the model burned while running the Index. Those do not change the Index score. They change whether 61 is a bargain. Astra's short-context list matches Fable at $10 / $50; Sol is cheaper per token at $4 / $20 with a promotional note. Finn's "price per task" point lives here: a higher $/1M model can still be cheaper per job if it uses far fewer tokens. That claim is testable in your harness; it is not what the Index number encodes. AA already publishes output tokens per Index task and cost per Index task on model pages — the efficiency metric behind that argument.
What AA grading is not
- Not ARC-AGI alone, and not a lab's SWE slide.
- Not "AGI" branding or the "Critical" cyber rating (that is an OpenAI Preparedness Framework self-assessment, not a third-party score).
- Not a substitute for trying Astra on your agent stack for computer use, gated cyber tools, or token efficiency.
When our Models page says Astra is 61, read it as: on this independent composite, at max effort, as of our last refresh (3 Sep 2026 — see How we source). When Finn says it destroys every benchmark, read it as: on the launch charts OpenAI (or a leaked deck) chose to show.
A practical reading checklist
Before you let a launch video re-rank your stack, run five questions:
- Who built the chart? Lab blog, "leaked" deck, third-party eval, or creator anecdote?
- Is it one bench or a portfolio? A single ARC number can swing by tens of points depending on version and scaffolding — the 98.6% vs 7% pairing above is the example. A composite is harder to game with one outlier.
- Are the competitors' settings matched? Max vs default, tools on vs off, same Terminal-Bench / SWE definition.
- Does price-per-task get equal time? Finn correctly separates list price from tokens-per-task. That claim is interesting, testable, and separate from "wins every bench."
- What does the independent board say today? For us, that is AA. Bookmark the model page, not just the thumbnail.
Hands On rule of thumb: treat lab launch benches as product claims until a public eval reproduces them. Product claims still matter — especially for agent workflows — but they are not the same as a leaderboard.
What to do with Astra this week
If you are choosing models for real work:
- Put GPT-6 Astra on your shortlist for hard multi-step agent jobs and gated cyber tooling. OpenAI is clearly positioning it as the GPT-6 flagship.
- Do not assume it already tops every public quality ranking. On AA Index, it does not.
- Compare price per completed task in your harness, not only $/1M tokens. Finn's video is useful here even when the scoreboard is lab-picked.
- Keep Fable 5.1 / Mythos 5.1 in the bakeoff until your own evals say otherwise. The AA gap is small enough that workflow fit will decide more than a YouTube chart.
And when the next creator says a model "blows everything out," pause the autoplay long enough to ask which everything, on whose sheet.
Sources
- Alex Finn — ChatGPT 6 Astra has released. AGI is here. (benchmark segment at 1:22)
- Artificial Analysis — GPT-6 Astra (max)
- Artificial Analysis — GPT-6 Astra (low)
- Artificial Analysis Intelligence Index v4.1.1 — methodology
- OpenAI — Updating our Preparedness Framework
- OpenAI — GPT-6 Astra model docs
- OpenAI — API pricing
- The Aigentic — Models
- The Aigentic — How we source
