OpenAI · launched 4 Sept 2026
GPT-6 Astra
OpenAI's flagship GPT-6 model, positioned above GPT-6 Sol and Luna for the most demanding professional work, coding and computer use.
Read the official announcement- Vendor
- OpenAI
- Launch date
- 4 Sept 2026
- Input price
- $10 / 1M tokens
- Output price
- $50 / 1M tokens
- Cached input
- $1.00 / 1M tokens
- Batch discount
- 50%
- Context window
- 1.05M tokens
- Max output
- 128K tokens
- Input types
- Text, Image, File
- OpenRouter ID
- openai/gpt-6-astra
- Terminal-Bench 4.0high effort · $7.21 per task · reported by OpenAI57.9%
- AutomationBench 1.0.6max effort · $1.73 per task · reported by OpenAI41.4%
- Agents' Last Exam V1max effort · $6.23 per task · reported by OpenAI59.3%
- Factual error rate(lower is better)high effort · $0.48 per task · reported by OpenAI3.9%
- FrontierCode 1.1max effort · $4.59 per task · reported by OpenAI53.3%
- DeepSWE 1.1xhigh effort · $4.43 per task · reported by OpenAI74.1%
- OSWorld 2.0max effort · $9.07 per task · reported by OpenAI73.5%
- Terminal-Bench 4.0high effort · $7.21 per attempt · reported by Anthropic57.9%
- FrontierCode v1.1max effort · $4.36 per task · reported by Anthropic53.26%
- GDPval-AA v2.1max effort · $4.53 per task · reported by Anthropic1542 Elo
- AutomationBenchmax effort · $1.77 per task · reported by Anthropic41.4%
Price compared with other models
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M | vs GPT-6 Astra |
|---|---|---|---|---|---|---|
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 99% cheaper |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 | 98% cheaper |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| Claude Sonnet 5.5Anthropic | 28 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| GPT-6 SolOpenAI | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 80% cheaper |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 | 60% cheaper |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 | 60% cheaper |
| Claude Opus 5Anthropic | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 | 50% cheaper |
| Claude Fable 5Anthropic | 9 Jun 2026 | $10 | $50 | $1.00 | $20 | Same price |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 | Same price |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 | Baseline |
USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token).
Official benchmark results
Reported by OpenAI on 22 Sept 2026 · exact values
| Benchmark | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | GPT-5.6 Sol | GPT-5.6 Luna |
|---|---|---|---|---|---|
| Coding deception rateAlignment · max effort · lower is better | 0.5% | 1.3% | 2.8% | 10.4% | 9.5% |
- Bold marks the best reported score in each row.
- OpenAI's internal evaluation uses tasks deliberately chosen to provoke dishonesty, so deception is much rarer in typical use.
- The announcement also covers broken search, reviewer bypass, warning circumvention and unauthorized interaction. See OpenAI's system card for those results.
Reported by Anthropic on 22 Sept 2026 · exact values
| Benchmark | Claude Opus 5.5 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (main)Agentic coding | 54.4% | 50.3% | 48% | 53.3% | 47.5% |
| CursorBench 4.0Agentic coding | 57.8% | 51.8% | 46.6% | – | 41.7% |
| GDPval-AA v2.1Knowledge work | 1846 Elo | 1735 Elo | 1708 Elo | 1542 Elo | 1588 Elo |
| AutomationBenchBusiness workflows | 40% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 67.7% | 65.6% | 63.6% | 57.2% | – |
| Terminal-Bench-Science 0.1Agentic scientific research | 58.7% | 52.6% | 29% | 64.6% | 22.4% |
| OSWorld 2.1Computer use · partial | 81.8% | 80.7% | 74% | – | – |
| ChartographyVisual chart recognition · with tools | 89% | 88.4% | 83.4% | – | – |
- Bold marks the best reported score in each row.
- Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
- Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
- A dash means the lab did not report a result for that model.
Accuracy against API cost · reported by OpenAI on 9 Sept 2026 · exact values
Tests agents on complex terminal-based tasks, including software engineering, system configuration and data analysis.
- GPT-6 Astra
- GPT-5.6 Sol
- Claude Fable 5.1
- Claude Fable 5
- Claude Opus 5
Show data table
| Model | Effort | Accuracy | Cost per task |
|---|---|---|---|
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Fable 5 | max | 44.5% | $22.2 |
| Claude Opus 5 | max | 52.6% | $18.4 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5 and Claude Opus 5 are shown at max effort only, as published.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance and HR.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 1.2% | $0.006 |
| GPT-6 Luna | medium | 9.4% | $0.016 |
| GPT-6 Luna | high | 14.5% | $0.021 |
| GPT-6 Luna | xhigh | 12.6% | $0.025 |
| GPT-6 Luna | max | 20.7% | $0.037 |
| GPT-5.6 Luna | low | 1.8% | $0.01 |
| GPT-5.6 Luna | medium | 4.3% | $0.02 |
| GPT-5.6 Luna | high | 9.1% | $0.05 |
| GPT-5.6 Luna | xhigh | 12.9% | $0.06 |
| GPT-5.6 Luna | max | 17% | $0.07 |
| GPT-6 Sol | low | 21.2% | $0.19 |
| GPT-6 Sol | medium | 26.9% | $0.21 |
| GPT-6 Sol | high | 31.2% | $0.24 |
| GPT-6 Sol | xhigh | 33.2% | $0.27 |
| GPT-6 Sol | max | 32% | $0.34 |
| GPT-5.6 Sol | low | 11.7% | $0.31 |
| GPT-5.6 Sol | medium | 19.6% | $0.42 |
| GPT-5.6 Sol | high | 24.8% | $0.47 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.54 |
| GPT-5.6 Sol | max | 28.8% | $0.67 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.27 |
| GPT-6 Astra | high | 37.1% | $1.44 |
| GPT-6 Astra | xhigh | 39% | $1.50 |
| GPT-6 Astra | max | 41.4% | $1.73 |
| Claude Opus 5 | low | 20.4% | $1.64 |
| Claude Opus 5 | medium | 23.9% | $2.22 |
| Claude Opus 5 | high | 20.5% | $2.27 |
| Claude Opus 5 | xhigh | 25.3% | $2.71 |
| Claude Opus 5 | max | 26.9% | $3.05 |
| Claude Fable 5.1 | max | 31.4% | $2.45 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5.1 ran with Claude Opus 5 as a fallback. OpenAI notes its cost omits the fallback runs (about 40% of tasks), so its real cost is higher.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries of professional computer work.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 36.3% | $0.025 |
| GPT-6 Luna | medium | 46.8% | $0.11 |
| GPT-6 Luna | high | 43.6% | $0.11 |
| GPT-6 Luna | xhigh | 47.9% | $0.11 |
| GPT-6 Luna | max | 50.9% | $0.15 |
| GPT-5.6 Luna | low | 31% | $0.18 |
| GPT-5.6 Luna | medium | 38.5% | $0.38 |
| GPT-5.6 Luna | high | 46.1% | $0.84 |
| GPT-5.6 Luna | xhigh | 49.4% | $1.55 |
| GPT-5.6 Luna | max | 50.4% | $2.57 |
| GPT-6 Sol | low | 48.7% | $0.86 |
| GPT-6 Sol | medium | 53.1% | $1.27 |
| GPT-6 Sol | high | 52.6% | $1.53 |
| GPT-6 Sol | xhigh | 55.4% | $1.67 |
| GPT-6 Sol | max | 56.4% | $2.93 |
| GPT-5.6 Sol | low | 45.1% | $1.63 |
| GPT-5.6 Sol | medium | 52.1% | $3.41 |
| GPT-5.6 Sol | high | 52.4% | $3.79 |
| GPT-5.6 Sol | xhigh | 53.6% | $5.08 |
| GPT-5.6 Sol | max | 52.8% | $7.13 |
| GPT-6 Astra | low | 53.4% | $3.07 |
| GPT-6 Astra | medium | 57.6% | $4.10 |
| GPT-6 Astra | high | 57.8% | $4.64 |
| GPT-6 Astra | xhigh | 58.3% | $5.40 |
| GPT-6 Astra | max | 59.3% | $6.23 |
| Claude Opus 5 | low | 51.9% | $4.06 |
| Claude Opus 5 | medium | 53% | $5.29 |
| Claude Opus 5 | high | 55.9% | $7.29 |
| Claude Opus 5 | xhigh | 55.5% | $10 |
| Claude Opus 5 | max | 52.7% | $9.76 |
| Claude Fable 5 | xhigh | 48.7% | $28.6 |
| Claude Fable 5 | adaptive | 41.3% | $15.2 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5 ran with Claude Opus 4.8 as a fallback.
Answers with any factual error against cost · lower is better · reported by OpenAI on 22 Sept 2026 · exact values
OpenAI's internal evaluation on de-identified ChatGPT conversations where users had flagged a factual error from a prior model.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
Show data table
| Model | Effort | Error rate | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 27.7% | $0.0024 |
| GPT-6 Luna | medium | 17.5% | $0.0033 |
| GPT-6 Luna | high | 12.5% | $0.0045 |
| GPT-6 Luna | xhigh | 10.2% | $0.0062 |
| GPT-6 Luna | max | 7.6% | $0.012 |
| GPT-5.6 Luna | low | 36.8% | $0.0045 |
| GPT-5.6 Luna | medium | 24.7% | $0.0062 |
| GPT-5.6 Luna | high | 17.3% | $0.011 |
| GPT-5.6 Luna | xhigh | 12.2% | $0.015 |
| GPT-5.6 Luna | max | 12% | $0.022 |
| GPT-6 Sol | low | 11.4% | $0.05 |
| GPT-6 Sol | medium | 6.9% | $0.069 |
| GPT-6 Sol | high | 5.1% | $0.099 |
| GPT-6 Sol | xhigh | 4.5% | $0.13 |
| GPT-6 Sol | max | 4.6% | $0.18 |
| GPT-5.6 Sol | low | 19.4% | $0.10 |
| GPT-5.6 Sol | medium | 15% | $0.15 |
| GPT-5.6 Sol | high | 10.8% | $0.23 |
| GPT-5.6 Sol | xhigh | 8.4% | $0.39 |
| GPT-5.6 Sol | max | 8.5% | $0.87 |
| GPT-6 Astra | low | 6.3% | $0.24 |
| GPT-6 Astra | medium | 4.4% | $0.31 |
| GPT-6 Astra | high | 3.9% | $0.48 |
| GPT-6 Astra | xhigh | 4% | $0.60 |
| GPT-6 Astra | max | 3.9% | $0.79 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- These prompts were chosen because they caused errors before, so error rates are far higher than in typical use.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents write code graded on correctness and mergeability: test quality, scope discipline, code style and codebase standards.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 25.7% | $0.021 |
| GPT-6 Luna | medium | 35.5% | $0.053 |
| GPT-6 Luna | high | 37.3% | $0.067 |
| GPT-6 Luna | xhigh | 37.1% | $0.073 |
| GPT-6 Luna | max | 42.4% | $0.11 |
| GPT-5.6 Luna | low | 15.4% | $0.06 |
| GPT-5.6 Luna | medium | 25.7% | $0.13 |
| GPT-5.6 Luna | high | 35.9% | $0.23 |
| GPT-5.6 Luna | xhigh | 38.9% | $0.31 |
| GPT-5.6 Luna | max | 39.8% | $0.37 |
| GPT-6 Sol | low | 37.3% | $0.45 |
| GPT-6 Sol | medium | 45.9% | $0.80 |
| GPT-6 Sol | high | 47.7% | $1.08 |
| GPT-6 Sol | xhigh | 48.4% | $1.37 |
| GPT-6 Sol | max | 49.3% | $2.14 |
| GPT-5.6 Sol | low | 35.4% | $1.89 |
| GPT-5.6 Sol | medium | 39.9% | $2.69 |
| GPT-5.6 Sol | high | 45.1% | $3.48 |
| GPT-5.6 Sol | xhigh | 46.8% | $4.15 |
| GPT-5.6 Sol | max | 47.5% | $5.19 |
| GPT-6 Astra | low | 45.3% | $1.70 |
| GPT-6 Astra | medium | 48.8% | $2.43 |
| GPT-6 Astra | high | 50.9% | $3.01 |
| GPT-6 Astra | xhigh | 50.6% | $3.28 |
| GPT-6 Astra | max | 53.3% | $4.59 |
| Claude Opus 5 | low | 41.9% | $2.68 |
| Claude Opus 5 | medium | 53.4% | $4.31 |
| Claude Opus 5 | high | 48% | $7.24 |
| Claude Opus 5 | xhigh | 43.6% | $9.14 |
| Claude Opus 5 | max | 48% | $11.4 |
| Claude Fable 5.1 | low | 49.8% | $2.38 |
| Claude Fable 5.1 | medium | 50.9% | $3.28 |
| Claude Fable 5.1 | high | 50.3% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.7% | $9.27 |
| Claude Fable 5.1 | max | 50.3% | $12.8 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents solve original, long-horizon software engineering tasks in real codebases.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 2.4% | $0.0057 |
| GPT-6 Luna | medium | 44.5% | $0.052 |
| GPT-6 Luna | high | 59.3% | $0.084 |
| GPT-6 Luna | xhigh | 61.3% | $0.11 |
| GPT-6 Luna | max | 66.6% | $0.22 |
| GPT-5.6 Luna | low | 1.2% | $0.011 |
| GPT-5.6 Luna | medium | 9.3% | $0.031 |
| GPT-5.6 Luna | high | 42.4% | $0.13 |
| GPT-5.6 Luna | xhigh | 56.2% | $0.27 |
| GPT-5.6 Luna | max | 62.2% | $0.53 |
| GPT-6 Sol | low | 37.2% | $0.16 |
| GPT-6 Sol | medium | 56.6% | $0.38 |
| GPT-6 Sol | high | 65.3% | $0.64 |
| GPT-6 Sol | xhigh | 66.6% | $1.00 |
| GPT-6 Sol | max | 68.8% | $2.74 |
| GPT-5.6 Sol | low | 45.4% | $0.82 |
| GPT-5.6 Sol | medium | 61.1% | $1.42 |
| GPT-5.6 Sol | high | 69.4% | $2.66 |
| GPT-5.6 Sol | xhigh | 70.7% | $3.60 |
| GPT-5.6 Sol | max | 72.7% | $6.46 |
| GPT-6 Astra | low | 67% | $1.60 |
| GPT-6 Astra | medium | 72.8% | $3.08 |
| GPT-6 Astra | high | 73.2% | $3.92 |
| GPT-6 Astra | xhigh | 74.1% | $4.43 |
| GPT-6 Astra | max | 73.2% | $7.50 |
| Claude Opus 5 | low | 58.1% | $1.66 |
| Claude Opus 5 | medium | 68.9% | $3.29 |
| Claude Opus 5 | high | 72.8% | $6.08 |
| Claude Opus 5 | xhigh | 73.2% | $9.07 |
| Claude Opus 5 | max | 73.7% | $11.8 |
| Claude Fable 5 | low | 59.6% | $3.76 |
| Claude Fable 5 | medium | 65.4% | $6.09 |
| Claude Fable 5 | high | 68.6% | $9.18 |
| Claude Fable 5 | xhigh | 69.9% | $13.4 |
| Claude Fable 5 | max | 69.7% | $21.6 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- OpenAI used Claude Fable 5 scores because Fable 5.1 scores were unavailable.
Partial reward against cost · release v2026.08.08 · reported by OpenAI on 22 Sept 2026 · exact values
AI agents attempt long-horizon computer-use workflows covering everyday and professional tasks.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 8.3% | $0.03 |
| GPT-6 Luna | medium | 31.5% | $0.062 |
| GPT-6 Luna | high | 41.4% | $0.12 |
| GPT-6 Luna | xhigh | 46.7% | $0.16 |
| GPT-6 Luna | max | 52.7% | $0.27 |
| GPT-5.6 Luna | low | 12% | $0.019 |
| GPT-5.6 Luna | medium | 22.7% | $0.052 |
| GPT-5.6 Luna | high | 35.9% | $0.16 |
| GPT-5.6 Luna | xhigh | 48% | $0.36 |
| GPT-5.6 Luna | max | 52.7% | $0.49 |
| GPT-6 Sol | low | 43.9% | $0.97 |
| GPT-6 Sol | medium | 54% | $1.32 |
| GPT-6 Sol | high | 58.3% | $1.64 |
| GPT-6 Sol | xhigh | 60.5% | $2.21 |
| GPT-6 Sol | max | 64.4% | $3.25 |
| GPT-5.6 Sol | low | 29.8% | $0.91 |
| GPT-5.6 Sol | medium | 49.7% | $2.73 |
| GPT-5.6 Sol | high | 56.5% | $4.46 |
| GPT-5.6 Sol | xhigh | 60.9% | $5.93 |
| GPT-5.6 Sol | max | 66.2% | $7.71 |
| GPT-6 Astra | low | 62.2% | $2.55 |
| GPT-6 Astra | medium | 69.3% | $5.10 |
| GPT-6 Astra | high | 70% | $6.60 |
| GPT-6 Astra | xhigh | 71.3% | $7.17 |
| GPT-6 Astra | max | 73.5% | $9.07 |
| Claude Opus 5 | low | 55.2% | $9.88 |
| Claude Opus 5 | medium | 60.3% | $12.7 |
| Claude Opus 5 | high | 65.9% | $15.9 |
| Claude Opus 5 | xhigh | 70.1% | $23.9 |
| Claude Opus 5 | max | 70.2% | $24.1 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Opus 5 values match the official OSWorld 2.0 leaderboard (osworld-v2.xlang.ai).
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command line interface.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Opus 5 | low | 28.5% | $4.25 |
| Claude Opus 5 | medium | 41.2% | $7.00 |
| Claude Opus 5 | high | 47% | $10.6 |
| Claude Opus 5 | xhigh | 50.6% | $13.5 |
| Claude Opus 5 | max | 52.3% | $15.8 |
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.64% | $0.80 |
| Claude Opus 5.5 | high | 53.99% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.42% | $2.25 |
| Claude Opus 5.5 | max | 54.43% | $6.19 |
| Claude Fable 5.1 | low | 52.8% | $2.47 |
| Claude Fable 5.1 | medium | 50.91% | $3.28 |
| Claude Fable 5.1 | high | 50.34% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.73% | $9.27 |
| Claude Fable 5.1 | max | 50.28% | $12.8 |
| Claude Opus 5 | low | 41.95% | $2.64 |
| Claude Opus 5 | medium | 53.38% | $4.61 |
| Claude Opus 5 | high | 47.99% | $7.62 |
| Claude Opus 5 | xhigh | 43.65% | $8.99 |
| Claude Opus 5 | max | 48.04% | $12.3 |
| GPT-6 Astra | low | 45.27% | $1.59 |
| GPT-6 Astra | medium | 48.83% | $2.28 |
| GPT-6 Astra | high | 50.94% | $2.85 |
| GPT-6 Astra | xhigh | 50.62% | $3.10 |
| GPT-6 Astra | max | 53.26% | $4.36 |
| GPT-5.6 Sol | low | 35.44% | $1.75 |
| GPT-5.6 Sol | medium | 39.93% | $2.50 |
| GPT-5.6 Sol | high | 45.06% | $3.25 |
| GPT-5.6 Sol | xhigh | 46.84% | $3.88 |
| GPT-5.6 Sol | max | 47.49% | $4.85 |
- Exact values from the data published with Anthropic's announcement.
Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values
Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 1224 Elo | $0.21 |
| Claude Opus 5.5 | medium | 1576 Elo | $0.86 |
| Claude Opus 5.5 | high | 1692 Elo | $1.54 |
| Claude Opus 5.5 | xhigh | 1820 Elo | $4.21 |
| Claude Opus 5.5 | max | 1846 Elo | $8.92 |
| Claude Fable 5.1 | low | 1450 Elo | $1.41 |
| Claude Fable 5.1 | medium | 1536 Elo | $2.17 |
| Claude Fable 5.1 | high | 1617 Elo | $3.43 |
| Claude Fable 5.1 | xhigh | 1721 Elo | $7.09 |
| Claude Fable 5.1 | max | 1735 Elo | $9.59 |
| Claude Opus 5 | low | 1294 Elo | $0.58 |
| Claude Opus 5 | medium | 1476 Elo | $1.36 |
| Claude Opus 5 | high | 1581 Elo | $3.03 |
| Claude Opus 5 | xhigh | 1676 Elo | $4.96 |
| Claude Opus 5 | max | 1708 Elo | $6.76 |
| GPT-6 Astra | low | 1366 Elo | $0.85 |
| GPT-6 Astra | medium | 1468 Elo | $1.82 |
| GPT-6 Astra | high | 1485 Elo | $2.43 |
| GPT-6 Astra | xhigh | 1516 Elo | $3.04 |
| GPT-6 Astra | max | 1542 Elo | $4.53 |
| GPT-5.6 Sol | low | 1289 Elo | $0.27 |
| GPT-5.6 Sol | medium | 1403 Elo | $0.60 |
| GPT-5.6 Sol | high | 1480 Elo | $1.11 |
| GPT-5.6 Sol | xhigh | 1548 Elo | $1.70 |
| GPT-5.6 Sol | max | 1588 Elo | $2.81 |
- Exact values from the data published with Anthropic's announcement.
Pass rate against cost · reported by Anthropic on 22 Sept 2026 · exact values
Built by Zapier, tests whether an agent can carry out real business workflows across many connected apps.
- Claude Opus 5.5
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 23.3% | $0.50 |
| Claude Opus 5.5 | medium | 28.6% | $0.64 |
| Claude Opus 5.5 | high | 32% | $0.70 |
| Claude Opus 5.5 | xhigh | 34.4% | $0.86 |
| Claude Opus 5.5 | max | 40% | $1.37 |
| Claude Opus 5 | low | 20.4% | $0.75 |
| Claude Opus 5 | medium | 23.9% | $0.89 |
| Claude Opus 5 | high | 20.6% | $1.03 |
| Claude Opus 5 | xhigh | 25.3% | $1.15 |
| Claude Opus 5 | max | 26.9% | $1.27 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.28 |
| GPT-6 Astra | high | 37.1% | $1.45 |
| GPT-6 Astra | xhigh | 39% | $1.53 |
| GPT-6 Astra | max | 41.4% | $1.77 |
| GPT-5.6 Sol | low | 11.7% | $0.42 |
| GPT-5.6 Sol | medium | 19.6% | $0.57 |
| GPT-5.6 Sol | high | 24.8% | $0.64 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.74 |
| GPT-5.6 Sol | max | 28.8% | $0.91 |
- Exact values from the data published with Anthropic's announcement.
- Runs were performed by Zapier without fallback models, so safeguard interventions counted as failures.