AI Model Tracker
Recent model launches with their API prices and the benchmark results published by each lab, linked back to the official announcement.
Latest releases
Anthropic · launched 28 Sept 2026
The second model in the Claude 5.5 family. Priced the same as Sonnet 5, it runs more than 30% faster and needs fewer tokens per task.
- Input / 1M
- $2.00
- Output / 1M
- $10
- Context
- 1M
Anthropic · launched 22 Sept 2026
Anthropic's newest Opus model. Tokens cost 20% less than Opus 5 and cache reads 60% less, with output generated more than 30% faster.
- Input / 1M
- $4.00
- Output / 1M
- $20
- Context
- 1M
OpenAI · launched 22 Sept 2026
The lowest-cost model in the GPT-6 family, positioned below GPT-6 Sol and the flagship GPT-6 Astra.
- Input / 1M
- $0.10
- Output / 1M
- $0.50
- Context
- 1.05M
OpenAI · launched 22 Sept 2026
The mid-tier GPT-6 model, sitting between GPT-6 Luna and the flagship GPT-6 Astra.
- Input / 1M
- $2.00
- Output / 1M
- $10
- Context
- 1.05M
OpenAI · launched 4 Sept 2026
OpenAI's flagship GPT-6 model, positioned above GPT-6 Sol and Luna for the most demanding professional work, coding and computer use.
- Input / 1M
- $10
- Output / 1M
- $50
- Context
- 1.05M
Anthropic · launched 1 Sept 2026
Released alongside the access-limited Claude Mythos 5.1. Cache reads cost 75% less than Fable 5, which Anthropic estimates cuts typical costs by about 25%.
- Input / 1M
- $10
- Output / 1M
- $50
- Context
- 1M
API pricing
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M |
|---|---|---|---|---|---|
| Claude Sonnet 5.5Anthropic | 28 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 |
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 |
| GPT-6 SolOpenAI | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 |
| Claude Opus 5Anthropic | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 |
| Claude Fable 5Anthropic | 9 Jun 2026 | $10 | $50 | $1.00 | $20 |
Standard API prices in USD per million tokens, checked 2 Oct 2026 via OpenRouter. Blended assumes 3 input tokens for every output token. Batch requests are 50% cheaper for every model listed.
Official benchmark results
OpenAI: Introducing GPT-6 Sol and LunaRead announcement
Reported by OpenAI on 22 Sept 2026 · exact values
| Benchmark | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | GPT-5.6 Sol | GPT-5.6 Luna |
|---|---|---|---|---|---|
| Coding deception rateAlignment · max effort · lower is better | 0.5% | 1.3% | 2.8% | 10.4% | 9.5% |
- Bold marks the best reported score in each row.
- OpenAI's internal evaluation uses tasks deliberately chosen to provoke dishonesty, so deception is much rarer in typical use.
- The announcement also covers broken search, reviewer bypass, warning circumvention and unauthorized interaction. See OpenAI's system card for those results.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents are tested on end-to-end workflows using 47 tools across sales, marketing, operations, support, finance and HR.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 1.2% | $0.006 |
| GPT-6 Luna | medium | 9.4% | $0.016 |
| GPT-6 Luna | high | 14.5% | $0.021 |
| GPT-6 Luna | xhigh | 12.6% | $0.025 |
| GPT-6 Luna | max | 20.7% | $0.037 |
| GPT-5.6 Luna | low | 1.8% | $0.01 |
| GPT-5.6 Luna | medium | 4.3% | $0.02 |
| GPT-5.6 Luna | high | 9.1% | $0.05 |
| GPT-5.6 Luna | xhigh | 12.9% | $0.06 |
| GPT-5.6 Luna | max | 17% | $0.07 |
| GPT-6 Sol | low | 21.2% | $0.19 |
| GPT-6 Sol | medium | 26.9% | $0.21 |
| GPT-6 Sol | high | 31.2% | $0.24 |
| GPT-6 Sol | xhigh | 33.2% | $0.27 |
| GPT-6 Sol | max | 32% | $0.34 |
| GPT-5.6 Sol | low | 11.7% | $0.31 |
| GPT-5.6 Sol | medium | 19.6% | $0.42 |
| GPT-5.6 Sol | high | 24.8% | $0.47 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.54 |
| GPT-5.6 Sol | max | 28.8% | $0.67 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.27 |
| GPT-6 Astra | high | 37.1% | $1.44 |
| GPT-6 Astra | xhigh | 39% | $1.50 |
| GPT-6 Astra | max | 41.4% | $1.73 |
| Claude Opus 5 | low | 20.4% | $1.64 |
| Claude Opus 5 | medium | 23.9% | $2.22 |
| Claude Opus 5 | high | 20.5% | $2.27 |
| Claude Opus 5 | xhigh | 25.3% | $2.71 |
| Claude Opus 5 | max | 26.9% | $3.05 |
| Claude Fable 5.1 | max | 31.4% | $2.45 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5.1 ran with Claude Opus 5 as a fallback. OpenAI notes its cost omits the fallback runs (about 40% of tasks), so its real cost is higher.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents are evaluated on long-horizon, economically valuable tasks spanning 55 sub-industries of professional computer work.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 36.3% | $0.025 |
| GPT-6 Luna | medium | 46.8% | $0.11 |
| GPT-6 Luna | high | 43.6% | $0.11 |
| GPT-6 Luna | xhigh | 47.9% | $0.11 |
| GPT-6 Luna | max | 50.9% | $0.15 |
| GPT-5.6 Luna | low | 31% | $0.18 |
| GPT-5.6 Luna | medium | 38.5% | $0.38 |
| GPT-5.6 Luna | high | 46.1% | $0.84 |
| GPT-5.6 Luna | xhigh | 49.4% | $1.55 |
| GPT-5.6 Luna | max | 50.4% | $2.57 |
| GPT-6 Sol | low | 48.7% | $0.86 |
| GPT-6 Sol | medium | 53.1% | $1.27 |
| GPT-6 Sol | high | 52.6% | $1.53 |
| GPT-6 Sol | xhigh | 55.4% | $1.67 |
| GPT-6 Sol | max | 56.4% | $2.93 |
| GPT-5.6 Sol | low | 45.1% | $1.63 |
| GPT-5.6 Sol | medium | 52.1% | $3.41 |
| GPT-5.6 Sol | high | 52.4% | $3.79 |
| GPT-5.6 Sol | xhigh | 53.6% | $5.08 |
| GPT-5.6 Sol | max | 52.8% | $7.13 |
| GPT-6 Astra | low | 53.4% | $3.07 |
| GPT-6 Astra | medium | 57.6% | $4.10 |
| GPT-6 Astra | high | 57.8% | $4.64 |
| GPT-6 Astra | xhigh | 58.3% | $5.40 |
| GPT-6 Astra | max | 59.3% | $6.23 |
| Claude Opus 5 | low | 51.9% | $4.06 |
| Claude Opus 5 | medium | 53% | $5.29 |
| Claude Opus 5 | high | 55.9% | $7.29 |
| Claude Opus 5 | xhigh | 55.5% | $10 |
| Claude Opus 5 | max | 52.7% | $9.76 |
| Claude Fable 5 | xhigh | 48.7% | $28.6 |
| Claude Fable 5 | adaptive | 41.3% | $15.2 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5 ran with Claude Opus 4.8 as a fallback.
Answers with any factual error against cost · lower is better · reported by OpenAI on 22 Sept 2026 · exact values
OpenAI's internal evaluation on de-identified ChatGPT conversations where users had flagged a factual error from a prior model.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
Show data table
| Model | Effort | Error rate | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 27.7% | $0.0024 |
| GPT-6 Luna | medium | 17.5% | $0.0033 |
| GPT-6 Luna | high | 12.5% | $0.0045 |
| GPT-6 Luna | xhigh | 10.2% | $0.0062 |
| GPT-6 Luna | max | 7.6% | $0.012 |
| GPT-5.6 Luna | low | 36.8% | $0.0045 |
| GPT-5.6 Luna | medium | 24.7% | $0.0062 |
| GPT-5.6 Luna | high | 17.3% | $0.011 |
| GPT-5.6 Luna | xhigh | 12.2% | $0.015 |
| GPT-5.6 Luna | max | 12% | $0.022 |
| GPT-6 Sol | low | 11.4% | $0.05 |
| GPT-6 Sol | medium | 6.9% | $0.069 |
| GPT-6 Sol | high | 5.1% | $0.099 |
| GPT-6 Sol | xhigh | 4.5% | $0.13 |
| GPT-6 Sol | max | 4.6% | $0.18 |
| GPT-5.6 Sol | low | 19.4% | $0.10 |
| GPT-5.6 Sol | medium | 15% | $0.15 |
| GPT-5.6 Sol | high | 10.8% | $0.23 |
| GPT-5.6 Sol | xhigh | 8.4% | $0.39 |
| GPT-5.6 Sol | max | 8.5% | $0.87 |
| GPT-6 Astra | low | 6.3% | $0.24 |
| GPT-6 Astra | medium | 4.4% | $0.31 |
| GPT-6 Astra | high | 3.9% | $0.48 |
| GPT-6 Astra | xhigh | 4% | $0.60 |
| GPT-6 Astra | max | 3.9% | $0.79 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- These prompts were chosen because they caused errors before, so error rates are far higher than in typical use.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents write code graded on correctness and mergeability: test quality, scope discipline, code style and codebase standards.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5.1
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 25.7% | $0.021 |
| GPT-6 Luna | medium | 35.5% | $0.053 |
| GPT-6 Luna | high | 37.3% | $0.067 |
| GPT-6 Luna | xhigh | 37.1% | $0.073 |
| GPT-6 Luna | max | 42.4% | $0.11 |
| GPT-5.6 Luna | low | 15.4% | $0.06 |
| GPT-5.6 Luna | medium | 25.7% | $0.13 |
| GPT-5.6 Luna | high | 35.9% | $0.23 |
| GPT-5.6 Luna | xhigh | 38.9% | $0.31 |
| GPT-5.6 Luna | max | 39.8% | $0.37 |
| GPT-6 Sol | low | 37.3% | $0.45 |
| GPT-6 Sol | medium | 45.9% | $0.80 |
| GPT-6 Sol | high | 47.7% | $1.08 |
| GPT-6 Sol | xhigh | 48.4% | $1.37 |
| GPT-6 Sol | max | 49.3% | $2.14 |
| GPT-5.6 Sol | low | 35.4% | $1.89 |
| GPT-5.6 Sol | medium | 39.9% | $2.69 |
| GPT-5.6 Sol | high | 45.1% | $3.48 |
| GPT-5.6 Sol | xhigh | 46.8% | $4.15 |
| GPT-5.6 Sol | max | 47.5% | $5.19 |
| GPT-6 Astra | low | 45.3% | $1.70 |
| GPT-6 Astra | medium | 48.8% | $2.43 |
| GPT-6 Astra | high | 50.9% | $3.01 |
| GPT-6 Astra | xhigh | 50.6% | $3.28 |
| GPT-6 Astra | max | 53.3% | $4.59 |
| Claude Opus 5 | low | 41.9% | $2.68 |
| Claude Opus 5 | medium | 53.4% | $4.31 |
| Claude Opus 5 | high | 48% | $7.24 |
| Claude Opus 5 | xhigh | 43.6% | $9.14 |
| Claude Opus 5 | max | 48% | $11.4 |
| Claude Fable 5.1 | low | 49.8% | $2.38 |
| Claude Fable 5.1 | medium | 50.9% | $3.28 |
| Claude Fable 5.1 | high | 50.3% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.7% | $9.27 |
| Claude Fable 5.1 | max | 50.3% | $12.8 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
Score against cost · reported by OpenAI on 22 Sept 2026 · exact values
AI agents solve original, long-horizon software engineering tasks in real codebases.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 2.4% | $0.0057 |
| GPT-6 Luna | medium | 44.5% | $0.052 |
| GPT-6 Luna | high | 59.3% | $0.084 |
| GPT-6 Luna | xhigh | 61.3% | $0.11 |
| GPT-6 Luna | max | 66.6% | $0.22 |
| GPT-5.6 Luna | low | 1.2% | $0.011 |
| GPT-5.6 Luna | medium | 9.3% | $0.031 |
| GPT-5.6 Luna | high | 42.4% | $0.13 |
| GPT-5.6 Luna | xhigh | 56.2% | $0.27 |
| GPT-5.6 Luna | max | 62.2% | $0.53 |
| GPT-6 Sol | low | 37.2% | $0.16 |
| GPT-6 Sol | medium | 56.6% | $0.38 |
| GPT-6 Sol | high | 65.3% | $0.64 |
| GPT-6 Sol | xhigh | 66.6% | $1.00 |
| GPT-6 Sol | max | 68.8% | $2.74 |
| GPT-5.6 Sol | low | 45.4% | $0.82 |
| GPT-5.6 Sol | medium | 61.1% | $1.42 |
| GPT-5.6 Sol | high | 69.4% | $2.66 |
| GPT-5.6 Sol | xhigh | 70.7% | $3.60 |
| GPT-5.6 Sol | max | 72.7% | $6.46 |
| GPT-6 Astra | low | 67% | $1.60 |
| GPT-6 Astra | medium | 72.8% | $3.08 |
| GPT-6 Astra | high | 73.2% | $3.92 |
| GPT-6 Astra | xhigh | 74.1% | $4.43 |
| GPT-6 Astra | max | 73.2% | $7.50 |
| Claude Opus 5 | low | 58.1% | $1.66 |
| Claude Opus 5 | medium | 68.9% | $3.29 |
| Claude Opus 5 | high | 72.8% | $6.08 |
| Claude Opus 5 | xhigh | 73.2% | $9.07 |
| Claude Opus 5 | max | 73.7% | $11.8 |
| Claude Fable 5 | low | 59.6% | $3.76 |
| Claude Fable 5 | medium | 65.4% | $6.09 |
| Claude Fable 5 | high | 68.6% | $9.18 |
| Claude Fable 5 | xhigh | 69.9% | $13.4 |
| Claude Fable 5 | max | 69.7% | $21.6 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- OpenAI used Claude Fable 5 scores because Fable 5.1 scores were unavailable.
Partial reward against cost · release v2026.08.08 · reported by OpenAI on 22 Sept 2026 · exact values
AI agents attempt long-horizon computer-use workflows covering everyday and professional tasks.
- GPT-6 Luna
- GPT-5.6 Luna
- GPT-6 Sol
- GPT-5.6 Sol
- GPT-6 Astra
- Claude Opus 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| GPT-6 Luna | low | 8.3% | $0.03 |
| GPT-6 Luna | medium | 31.5% | $0.062 |
| GPT-6 Luna | high | 41.4% | $0.12 |
| GPT-6 Luna | xhigh | 46.7% | $0.16 |
| GPT-6 Luna | max | 52.7% | $0.27 |
| GPT-5.6 Luna | low | 12% | $0.019 |
| GPT-5.6 Luna | medium | 22.7% | $0.052 |
| GPT-5.6 Luna | high | 35.9% | $0.16 |
| GPT-5.6 Luna | xhigh | 48% | $0.36 |
| GPT-5.6 Luna | max | 52.7% | $0.49 |
| GPT-6 Sol | low | 43.9% | $0.97 |
| GPT-6 Sol | medium | 54% | $1.32 |
| GPT-6 Sol | high | 58.3% | $1.64 |
| GPT-6 Sol | xhigh | 60.5% | $2.21 |
| GPT-6 Sol | max | 64.4% | $3.25 |
| GPT-5.6 Sol | low | 29.8% | $0.91 |
| GPT-5.6 Sol | medium | 49.7% | $2.73 |
| GPT-5.6 Sol | high | 56.5% | $4.46 |
| GPT-5.6 Sol | xhigh | 60.9% | $5.93 |
| GPT-5.6 Sol | max | 66.2% | $7.71 |
| GPT-6 Astra | low | 62.2% | $2.55 |
| GPT-6 Astra | medium | 69.3% | $5.10 |
| GPT-6 Astra | high | 70% | $6.60 |
| GPT-6 Astra | xhigh | 71.3% | $7.17 |
| GPT-6 Astra | max | 73.5% | $9.07 |
| Claude Opus 5 | low | 55.2% | $9.88 |
| Claude Opus 5 | medium | 60.3% | $12.7 |
| Claude Opus 5 | high | 65.9% | $15.9 |
| Claude Opus 5 | xhigh | 70.1% | $23.9 |
| Claude Opus 5 | max | 70.2% | $24.1 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Opus 5 values match the official OSWorld 2.0 leaderboard (osworld-v2.xlang.ai).
Anthropic: Introducing Claude Opus 5.5Read announcement
Reported by Anthropic on 22 Sept 2026 · exact values
| Benchmark | Claude Opus 5.5 | Claude Fable 5.1 | Claude Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (main)Agentic coding | 54.4% | 50.3% | 48% | 53.3% | 47.5% |
| CursorBench 4.0Agentic coding | 57.8% | 51.8% | 46.6% | – | 41.7% |
| GDPval-AA v2.1Knowledge work | 1846 Elo | 1735 Elo | 1708 Elo | 1542 Elo | 1588 Elo |
| AutomationBenchBusiness workflows | 40% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 67.7% | 65.6% | 63.6% | 57.2% | – |
| Terminal-Bench-Science 0.1Agentic scientific research | 58.7% | 52.6% | 29% | 64.6% | 22.4% |
| OSWorld 2.1Computer use · partial | 81.8% | 80.7% | 74% | – | – |
| ChartographyVisual chart recognition · with tools | 89% | 88.4% | 83.4% | – | – |
- Bold marks the best reported score in each row.
- Claude Opus 5.5 results use adaptive thinking at max effort unless noted.
- Terminal-Bench 4.0 shows each model's highest score: Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort. GPT-6 Astra and GPT-5.6 Sol figures there are as reported by OpenAI.
- A dash means the lab did not report a result for that model.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command line interface.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Opus 5 | low | 28.5% | $4.25 |
| Claude Opus 5 | medium | 41.2% | $7.00 |
| Claude Opus 5 | high | 47% | $10.6 |
| Claude Opus 5 | xhigh | 50.6% | $13.5 |
| Claude Opus 5 | max | 52.3% | $15.8 |
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.64% | $0.80 |
| Claude Opus 5.5 | high | 53.99% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.42% | $2.25 |
| Claude Opus 5.5 | max | 54.43% | $6.19 |
| Claude Fable 5.1 | low | 52.8% | $2.47 |
| Claude Fable 5.1 | medium | 50.91% | $3.28 |
| Claude Fable 5.1 | high | 50.34% | $5.27 |
| Claude Fable 5.1 | xhigh | 48.73% | $9.27 |
| Claude Fable 5.1 | max | 50.28% | $12.8 |
| Claude Opus 5 | low | 41.95% | $2.64 |
| Claude Opus 5 | medium | 53.38% | $4.61 |
| Claude Opus 5 | high | 47.99% | $7.62 |
| Claude Opus 5 | xhigh | 43.65% | $8.99 |
| Claude Opus 5 | max | 48.04% | $12.3 |
| GPT-6 Astra | low | 45.27% | $1.59 |
| GPT-6 Astra | medium | 48.83% | $2.28 |
| GPT-6 Astra | high | 50.94% | $2.85 |
| GPT-6 Astra | xhigh | 50.62% | $3.10 |
| GPT-6 Astra | max | 53.26% | $4.36 |
| GPT-5.6 Sol | low | 35.44% | $1.75 |
| GPT-5.6 Sol | medium | 39.93% | $2.50 |
| GPT-5.6 Sol | high | 45.06% | $3.25 |
| GPT-5.6 Sol | xhigh | 46.84% | $3.88 |
| GPT-5.6 Sol | max | 47.49% | $4.85 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 43.7% | $1.18 |
| Claude Opus 5.5 | medium | 52.5% | $2.90 |
| Claude Opus 5.5 | high | 56% | $3.97 |
| Claude Opus 5.5 | xhigh | 56% | $6.99 |
| Claude Opus 5.5 | max | 57.8% | $13.4 |
| Claude Fable 5.1 | low | 45.1% | $5.44 |
| Claude Fable 5.1 | medium | 46.8% | $7.05 |
| Claude Fable 5.1 | high | 49.2% | $9.08 |
| Claude Fable 5.1 | xhigh | 51.6% | $13 |
| Claude Fable 5.1 | max | 51.8% | $17.3 |
| Claude Opus 5 | low | 40.7% | $4.87 |
| Claude Opus 5 | medium | 43.3% | $6.94 |
| Claude Opus 5 | high | 44.7% | $9.00 |
| Claude Opus 5 | xhigh | 46.1% | $11.4 |
| Claude Opus 5 | max | 46.6% | $11.9 |
| GPT-5.6 Sol | low | 24.6% | $0.87 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 |
| GPT-5.6 Sol | high | 35.7% | $2.85 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 |
| GPT-5.6 Sol | max | 41.7% | $8.23 |
- Exact values from the data published with Anthropic's announcement.
Elo against cost · reported by Anthropic on 22 Sept 2026 · exact values
Artificial Analysis's evaluation of agents on real-world professional work across 44 occupations.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 1224 Elo | $0.21 |
| Claude Opus 5.5 | medium | 1576 Elo | $0.86 |
| Claude Opus 5.5 | high | 1692 Elo | $1.54 |
| Claude Opus 5.5 | xhigh | 1820 Elo | $4.21 |
| Claude Opus 5.5 | max | 1846 Elo | $8.92 |
| Claude Fable 5.1 | low | 1450 Elo | $1.41 |
| Claude Fable 5.1 | medium | 1536 Elo | $2.17 |
| Claude Fable 5.1 | high | 1617 Elo | $3.43 |
| Claude Fable 5.1 | xhigh | 1721 Elo | $7.09 |
| Claude Fable 5.1 | max | 1735 Elo | $9.59 |
| Claude Opus 5 | low | 1294 Elo | $0.58 |
| Claude Opus 5 | medium | 1476 Elo | $1.36 |
| Claude Opus 5 | high | 1581 Elo | $3.03 |
| Claude Opus 5 | xhigh | 1676 Elo | $4.96 |
| Claude Opus 5 | max | 1708 Elo | $6.76 |
| GPT-6 Astra | low | 1366 Elo | $0.85 |
| GPT-6 Astra | medium | 1468 Elo | $1.82 |
| GPT-6 Astra | high | 1485 Elo | $2.43 |
| GPT-6 Astra | xhigh | 1516 Elo | $3.04 |
| GPT-6 Astra | max | 1542 Elo | $4.53 |
| GPT-5.6 Sol | low | 1289 Elo | $0.27 |
| GPT-5.6 Sol | medium | 1403 Elo | $0.60 |
| GPT-5.6 Sol | high | 1480 Elo | $1.11 |
| GPT-5.6 Sol | xhigh | 1548 Elo | $1.70 |
| GPT-5.6 Sol | max | 1588 Elo | $2.81 |
- Exact values from the data published with Anthropic's announcement.
Pass rate against cost · reported by Anthropic on 22 Sept 2026 · exact values
Built by Zapier, tests whether an agent can carry out real business workflows across many connected apps.
- Claude Opus 5.5
- Claude Opus 5
- GPT-6 Astra
- GPT-5.6 Sol
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 23.3% | $0.50 |
| Claude Opus 5.5 | medium | 28.6% | $0.64 |
| Claude Opus 5.5 | high | 32% | $0.70 |
| Claude Opus 5.5 | xhigh | 34.4% | $0.86 |
| Claude Opus 5.5 | max | 40% | $1.37 |
| Claude Opus 5 | low | 20.4% | $0.75 |
| Claude Opus 5 | medium | 23.9% | $0.89 |
| Claude Opus 5 | high | 20.6% | $1.03 |
| Claude Opus 5 | xhigh | 25.3% | $1.15 |
| Claude Opus 5 | max | 26.9% | $1.27 |
| GPT-6 Astra | low | 30.3% | $1.08 |
| GPT-6 Astra | medium | 34.1% | $1.28 |
| GPT-6 Astra | high | 37.1% | $1.45 |
| GPT-6 Astra | xhigh | 39% | $1.53 |
| GPT-6 Astra | max | 41.4% | $1.77 |
| GPT-5.6 Sol | low | 11.7% | $0.42 |
| GPT-5.6 Sol | medium | 19.6% | $0.57 |
| GPT-5.6 Sol | high | 24.8% | $0.64 |
| GPT-5.6 Sol | xhigh | 26.3% | $0.74 |
| GPT-5.6 Sol | max | 28.8% | $0.91 |
- Exact values from the data published with Anthropic's announcement.
- Runs were performed by Zapier without fallback models, so safeguard interventions counted as failures.
Accuracy against cost · reported by Anthropic on 22 Sept 2026 · exact values
Perplexity's benchmark measuring agents on large data collection tasks.
- Claude Opus 5.5
- Claude Fable 5.1
- Claude Opus 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Opus 5.5 | low | 31.2% | $1.20 |
| Claude Opus 5.5 | medium | 62.8% | $11.2 |
| Claude Opus 5.5 | high | 67.3% | $16.9 |
| Claude Opus 5.5 | xhigh | 71.3% | $29.1 |
| Claude Opus 5.5 | max | 72.3% | $37.9 |
| Claude Fable 5.1 | low | 63.3% | $23.3 |
| Claude Fable 5.1 | medium | 64.5% | $27.4 |
| Claude Fable 5.1 | high | 66.7% | $33.6 |
| Claude Fable 5.1 | xhigh | 67.7% | $42.8 |
| Claude Fable 5.1 | max | 68.7% | $49 |
| Claude Opus 5 | low | 50.5% | $10.7 |
| Claude Opus 5 | medium | 58.1% | $24 |
| Claude Opus 5 | high | 64.6% | $43.5 |
| Claude Opus 5 | xhigh | 67% | $53.8 |
| Claude Opus 5 | max | 67.2% | $61.6 |
- Exact values from the data published with Anthropic's announcement.
- Claude models ran with offline web search and fetch tools and a 980K-token task budget. This differs from Perplexity's published setup, so scores are not comparable with Perplexity's leaderboard.
OpenAI: GPT-6 Astra: The next generation in intelligence for workRead announcement
Accuracy against API cost · reported by OpenAI on 9 Sept 2026 · exact values
Tests agents on complex terminal-based tasks, including software engineering, system configuration and data analysis.
- GPT-6 Astra
- GPT-5.6 Sol
- Claude Fable 5.1
- Claude Fable 5
- Claude Opus 5
Show data table
| Model | Effort | Accuracy | Cost per task |
|---|---|---|---|
| GPT-6 Astra | low | 49.7% | $4.95 |
| GPT-6 Astra | medium | 53.9% | $6.15 |
| GPT-6 Astra | high | 57.9% | $7.21 |
| GPT-6 Astra | xhigh | 57.6% | $7.48 |
| GPT-6 Astra | max | 56.7% | $10.3 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
| Claude Fable 5.1 | low | 40.2% | $5.70 |
| Claude Fable 5.1 | medium | 43.4% | $7.80 |
| Claude Fable 5.1 | high | 49.4% | $10.5 |
| Claude Fable 5.1 | xhigh | 51.3% | $15.8 |
| Claude Fable 5.1 | max | 55.8% | $19.5 |
| Claude Fable 5 | max | 44.5% | $22.2 |
| Claude Opus 5 | max | 52.6% | $18.4 |
- Exact values from the data labels on OpenAI's charts. OpenAI took competitor scores from publicly available reports.
- Claude Fable 5 and Claude Opus 5 are shown at max effort only, as published.
Anthropic: Introducing Claude Sonnet 5.5Read announcement
Reported by Anthropic on 28 Sept 2026 · exact values
| Benchmark | Claude Sonnet 5.5 | Claude Sonnet 5 | Claude Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0Agentic coding | 70.6% | 10.3% | 66.4% | – |
| FrontierCode 1.1 (main)Agentic coding · Sonnet 5.5 at max effort, 52.1% at xhigh | 46.2% | 42.4% | 54.4% | 49.3% |
| CursorBench 4.0Agentic coding | 55.5% | 34.1% | 57.8% | – |
| GDPval-AA v2.1Knowledge work | 1844 Elo | 1449 Elo | 1846 Elo | 1487 Elo |
| AA-Briefcase v1.1Knowledge work | 1811 Elo | 1359 Elo | 1822 Elo | 1483 Elo |
| Humanity's Last ExamMultidisciplinary reasoning · with tools | 64.5% | 54.9% | 67.7% | – |
| OSWorld 2.1Computer use · partial | 80.1% | 57% | 81.8% | – |
| ChartographyVisual chart recognition · no tools | 61.6% | 15.6% | 64.4% | 53.6% |
- Bold marks the best reported score in each row.
- Claude Opus 5.5's Terminal-Bench 4.0 score is at xhigh effort, its highest.
- OpenAI recently fixed a bug affecting GPT-6 Sol's image understanding; its GDPval-AA, AA-Briefcase and Chartography scores may predate the fix.
- A dash means the lab did not report a result for that model.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Measures how well a model completes complex, multi-step professional tasks within a command-line interface.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 20% | $0.76 |
| Claude Sonnet 5.5 | medium | 28.8% | $0.83 |
| Claude Sonnet 5.5 | high | 43% | $1.94 |
| Claude Sonnet 5.5 | xhigh | 61.5% | $5.30 |
| Claude Sonnet 5.5 | max | 70.6% | $12.5 |
| Claude Opus 5.5 | low | 38.5% | $1.29 |
| Claude Opus 5.5 | medium | 57.6% | $2.94 |
| Claude Opus 5.5 | high | 64.2% | $3.88 |
| Claude Opus 5.5 | xhigh | 66.4% | $7.35 |
| Claude Opus 5.5 | max | 64.8% | $11.2 |
| Claude Sonnet 5 | low | 3.2% | $3.78 |
| Claude Sonnet 5 | medium | 4.4% | $4.83 |
| Claude Sonnet 5 | high | 4.5% | $8.20 |
| Claude Sonnet 5 | xhigh | 7% | $9.95 |
| Claude Sonnet 5 | max | 10.3% | $11.6 |
| GPT-5.6 Sol | low | 7.9% | $1.46 |
| GPT-5.6 Sol | medium | 20.9% | $2.69 |
| GPT-5.6 Sol | high | 26.1% | $4.12 |
| GPT-5.6 Sol | xhigh | 28.5% | $5.39 |
| GPT-5.6 Sol | max | 37.3% | $7.89 |
- Exact values from the data published with Anthropic's announcement.
- OpenAI did not report GPT-6 Sol on this benchmark, so Anthropic shows GPT-5.6 Sol.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Measures whether an agent's code changes would be merged.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 29.3% | $0.19 |
| Claude Sonnet 5.5 | medium | 36.5% | $0.24 |
| Claude Sonnet 5.5 | high | 49.4% | $0.42 |
| Claude Sonnet 5.5 | xhigh | 52.1% | $1.59 |
| Claude Sonnet 5.5 | max | 46.2% | $20.8 |
| Claude Opus 5.5 | low | 47.3% | $0.40 |
| Claude Opus 5.5 | medium | 54.6% | $0.80 |
| Claude Opus 5.5 | high | 54% | $1.09 |
| Claude Opus 5.5 | xhigh | 51.4% | $2.25 |
| Claude Opus 5.5 | max | 54.4% | $6.19 |
| Claude Sonnet 5 | low | 28.7% | $2.39 |
| Claude Sonnet 5 | medium | 35.2% | $3.81 |
| Claude Sonnet 5 | high | 39.4% | $6.10 |
| Claude Sonnet 5 | xhigh | 42.7% | $10.1 |
| Claude Sonnet 5 | max | 42.4% | $17.1 |
| GPT-6 Sol | low | 37.3% | $0.43 |
| GPT-6 Sol | medium | 45.9% | $0.77 |
| GPT-6 Sol | high | 47.7% | $1.04 |
| GPT-6 Sol | xhigh | 48.4% | $1.32 |
| GPT-6 Sol | max | 49.3% | $2.07 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 28 Sept 2026 · exact values
Evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-5.6 Sol
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 35.8% | $0.50 |
| Claude Sonnet 5.5 | medium | 39.2% | $0.70 |
| Claude Sonnet 5.5 | high | 47.8% | $1.67 |
| Claude Sonnet 5.5 | xhigh | 53.1% | $3.88 |
| Claude Sonnet 5.5 | max | 55.5% | $9.67 |
| Claude Opus 5.5 | low | 43.7% | $1.17 |
| Claude Opus 5.5 | medium | 52.5% | $2.91 |
| Claude Opus 5.5 | high | 56% | $3.97 |
| Claude Opus 5.5 | xhigh | 56% | $6.98 |
| Claude Opus 5.5 | max | 57.8% | $13.4 |
| Claude Sonnet 5 | low | 24.1% | $1.39 |
| Claude Sonnet 5 | medium | 28% | $2.31 |
| Claude Sonnet 5 | high | 30.8% | $3.48 |
| Claude Sonnet 5 | xhigh | 32% | $4.55 |
| Claude Sonnet 5 | max | 34.1% | $7.17 |
| GPT-5.6 Sol | low | 24.6% | $0.87 |
| GPT-5.6 Sol | medium | 31.1% | $1.77 |
| GPT-5.6 Sol | high | 35.7% | $2.85 |
| GPT-5.6 Sol | xhigh | 37.7% | $4.40 |
| GPT-5.6 Sol | max | 41.7% | $8.23 |
- Exact values from the data published with Anthropic's announcement.
- CursorBench does not report GPT-6 Sol publicly, so Anthropic shows GPT-5.6 Sol.
Elo against cost · reported by Anthropic on 28 Sept 2026 · exact values
Artificial Analysis's benchmark of long-horizon knowledge work.
- Claude Sonnet 5.5
- Claude Opus 5.5
- Claude Sonnet 5
- GPT-6 Sol
Show data table
| Model | Effort | Elo | Cost per task |
|---|---|---|---|
| Claude Sonnet 5.5 | low | 1264 Elo | $0.87 |
| Claude Sonnet 5.5 | medium | 1461 Elo | $1.64 |
| Claude Sonnet 5.5 | high | 1634 Elo | $3.95 |
| Claude Sonnet 5.5 | xhigh | 1746 Elo | $9.63 |
| Claude Sonnet 5.5 | max | 1811 Elo | $29.2 |
| Claude Opus 5.5 | low | 1285 Elo | $1.15 |
| Claude Opus 5.5 | medium | 1642 Elo | $4.40 |
| Claude Opus 5.5 | high | 1705 Elo | $6.27 |
| Claude Opus 5.5 | xhigh | 1780 Elo | $12.3 |
| Claude Opus 5.5 | max | 1822 Elo | $21 |
| Claude Sonnet 5 | low | 923 Elo | $0.82 |
| Claude Sonnet 5 | medium | 1056 Elo | $1.73 |
| Claude Sonnet 5 | high | 1177 Elo | $3.76 |
| Claude Sonnet 5 | xhigh | 1274 Elo | $7.56 |
| Claude Sonnet 5 | max | 1359 Elo | $14.4 |
| GPT-6 Sol | low | 905 Elo | $0.12 |
| GPT-6 Sol | medium | 1142 Elo | $0.34 |
| GPT-6 Sol | high | 1289 Elo | $0.63 |
| GPT-6 Sol | xhigh | 1364 Elo | $1.19 |
| GPT-6 Sol | max | 1483 Elo | $2.67 |
- Exact values from the data published with Anthropic's announcement.
- Artificial Analysis ran Sonnet 5.5 on a pre-release deployment with a since-fixed structured output bug, which may slightly understate its score.
Anthropic: Introducing Claude Fable 5.1 and Claude Mythos 5.1Read announcement
Reported by Anthropic on 1 Sept 2026 · exact values
| Benchmark | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1Agentic scientific research | 52.6% | 24.7% | 29% | 22.4% |
| Terminal-Bench 4.0Agentic coding | 55.8% | 42% | 52.3% | 37.3% |
| GDPval-AA v2Knowledge work | 1853 Elo | 1723 Elo | 1824 Elo | 1711 Elo |
| OSWorld 2.0Computer use · partial | 77.9% | 72.9% | 75.4% | – |
| OSWorld 2.0 (strict)Computer use · strict | 41.7% | 36.1% | 39.6% | – |
| Humanity's Last Exam (no tools)Multidisciplinary reasoning · no tools | 60.9% | 57.8% | 56.6% | – |
| Humanity's Last Exam (with tools)Multidisciplinary reasoning · with tools | 65% | 63.8% | 63.6% | – |
| AutomationBenchBusiness workflows | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0Agentic coding | 73.4% | 70.5% | 70% | 67.2% |
- Bold marks the best reported score in each row.
- Fable 5.1 was evaluated with production safeguards enabled. Where they intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 and Fable 5 scored zero on AutomationBench.
- Claude Mythos 5.1, the same underlying model with access limited to vetted users, scores 60.9% on Terminal-Bench 4.0. Mythos has no public API price, so it is not tracked here.
- A dash means the lab did not report a result for that model.
Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values
Terminal-based agentic scientific research tasks, in Anthropic's agentic scientific research category.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 26.3% | $11.1 |
| Claude Fable 5.1 | medium | 35.7% | $14.9 |
| Claude Fable 5.1 | high | 40% | $20.3 |
| Claude Fable 5.1 | xhigh | 49.5% | $31.8 |
| Claude Fable 5.1 | max | 52.6% | $37.9 |
| Claude Fable 5 | low | 12.3% | $17.1 |
| Claude Fable 5 | medium | 21.4% | $25 |
| Claude Fable 5 | high | 25% | $34.3 |
| Claude Fable 5 | xhigh | 23.4% | $36 |
| Claude Fable 5 | max | 24.7% | $44.1 |
- Exact values from the data published with Anthropic's announcement.
- Standard error is about 3.5 to 4.5 points per model.
Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values
Expert-level questions across many disciplines, answered with access to tools.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 60% | $0.52 |
| Claude Fable 5.1 | medium | 62.96% | $0.67 |
| Claude Fable 5.1 | high | 64.76% | $1.05 |
| Claude Fable 5.1 | xhigh | 65.08% | $2.28 |
| Claude Fable 5.1 | max | 65% | $3.20 |
| Claude Fable 5 | low | 59.64% | $0.61 |
| Claude Fable 5 | medium | 61.4% | $1.01 |
| Claude Fable 5 | high | 63.16% | $1.42 |
| Claude Fable 5 | xhigh | 63.52% | $1.92 |
| Claude Fable 5 | max | 63.8% | $3.44 |
- Exact values from the data published with Anthropic's announcement.
Pass rate against cost · reported by Anthropic on 1 Sept 2026 · exact values
Expert-level questions across many disciplines, answered without tools.
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Pass rate | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 53.16% | $0.30 |
| Claude Fable 5.1 | medium | 55.92% | $0.46 |
| Claude Fable 5.1 | high | 57.96% | $0.75 |
| Claude Fable 5.1 | xhigh | 60.36% | $1.53 |
| Claude Fable 5.1 | max | 60.92% | $2.23 |
| Claude Fable 5 | low | 50.6% | $0.17 |
| Claude Fable 5 | medium | 55.88% | $0.40 |
| Claude Fable 5 | high | 56.88% | $0.62 |
| Claude Fable 5 | xhigh | 57.44% | $0.91 |
| Claude Fable 5 | max | 57.76% | $1.70 |
- Exact values from the data published with Anthropic's announcement.
Accuracy against cost · reported by Anthropic on 1 Sept 2026 · exact values
Evaluates coding agents on tasks taken from real Cursor sessions (earlier release than CursorBench 4.0).
- Claude Fable 5.1
- Claude Fable 5
Show data table
| Model | Effort | Score | Cost per task |
|---|---|---|---|
| Claude Fable 5.1 | low | 66.2% | $2.90 |
| Claude Fable 5.1 | medium | 68% | $3.53 |
| Claude Fable 5.1 | high | 69.4% | $4.80 |
| Claude Fable 5.1 | xhigh | 72.8% | $6.96 |
| Claude Fable 5.1 | max | 73.4% | $9.64 |
| Claude Fable 5 | low | 62.1% | $4.46 |
| Claude Fable 5 | medium | 65.2% | $6.80 |
| Claude Fable 5 | high | 66.5% | $8.77 |
| Claude Fable 5 | xhigh | 68.4% | $11.7 |
| Claude Fable 5 | max | 70.5% | $17.3 |
- Exact values from the data published with Anthropic's announcement.