AI Octane Index · BEP Research
Model prices are quoted per token, but you pay per job that comes back right — and a model that needs three tries costs three times. The same basket of work, run through the whole board, spans 273× in what it costs to get finished.
weighted cost per 10,000 successful attempts in this benchmark, including the cost of failed attempts. Answered 92 of 96 correctly (90–98% at 95% confidence).
the price of the pick, for 4 more points of observed success (96/96). This small sample does not establish the size of an accuracy premium. Overlapping intervals alone are not a significance test; a larger paired evaluation is needed.
Something on this board is cheaper and more accurate than each of them. The worst, GPT-5.6 Sol Pro, costs 273× the pick.
Tap or hover any dot to see which model it is.
Each dot is one model. Left is less accurate, up is more expensive per completed task, so the best place to be is bottom right. Anything inside the shaded region is beaten on both counts at once by something else on this board.
| Industry | Lowest cost above the test threshold | Cost per 10,000 completed | Correct |
|---|---|---|---|
| Banking & Finance | gpt-oss 120B | $0.38 | 18/18 |
| Customer Service | DeepSeek V4 Flash | $0.31 | 15/15 |
| Legal | gpt-oss 120B | $1.18 | 12/15 |
| Medical | DeepSeek V4 Flash | $0.28 | 27/27 |
| Software Engineering | DeepSeek V3.2 | $1.00 | 21/21 |
In 3 of 5 industries the cheapest deployable model is not the board's overall pick. Each industry is a slice of the same basket, so these rest on 15–27 attempts against the 96 behind the board above — a steer, not a ranking.
Calculate the interest accrued. LOAN TERMS: - Principal: $1,000,000 - Rate: 6.00% per annum - Day-count convention: Actual/360 - Accrual period: 45 days Respond with the interest in dollars, rounded to two decimal places, as a number only — no commas, no currency symbol.
7500.00Graded by exact comparison against that answer, so a right number wrapped in an explanation is scored wrong — the prompt asked for a number. Every model on the board gets this same prompt at temperature 0. See all 32 tasks.
Every prompt states an output format — “as a number only” — and the grader compares exactly, so a right answer wrapped in an explanation scores zero. That is a real failure if code reads the output, and no failure at all if a person does. 7 of 33 models are affected; the worst, Llama 4 Maverick, would score 81/96 instead of 56/96 — 25 of its answers were right and rejected on presentation. The ranking above uses the strict score; decide which column is yours.
| Model | Correct | Observed success | Right answer, any format | Cost per 10,000 completed | vs best | Verdict |
|---|---|---|---|---|---|---|
| gpt-oss 120B | 92/96 | 96% 90–98 | same | $0.70 | 1.0× | on the frontier |
| DeepSeek V3.2 | 83/96 | 86% 78–92 | same | $0.73 | 1.0× | beaten on both |
| GPT-5.6 Luna | 94/96 | 98% 93–99 | same | $1.64 | 2.3× | on the frontier |
| Gemini 2.5 Flash | 84/96 | 88% 79–93 | same | $3.22 | 4.6× | beaten on both |
| Llama 4 Maverick | 56/96 | 58% 48–68 | 81/96 +25 | $3.36 | 4.8× | format, not accuracy |
| GPT-5.4 mini | 79/96 | 82% 73–89 | same | $4.21 | 6.0× | beaten on both |
| Claude Haiku 4.5 | 75/96 | 78% 69–85 | 81/96 +6 | $8.25 | 11.7× | format, not accuracy |
| DeepSeek V4 Flash | 85/96 | 89% 81–93 | same | $8.51 | 12.1× | beaten on both |
| GPT-5.6 Luna Pro | 96/96 | 100% 96–100 | same | $9.44 | 13.4× | on the frontier |
| Gemini 3.7 Flash | 92/96 | 96% 90–98 | same | $9.88 | 14.0× | beaten on both |
| GPT-5.6 Sol | 93/96 | 97% 91–99 | same | $11.0 | 15.6× | beaten on both |
| GPT-5.6 Terra | 93/96 | 97% 91–99 | same | $13.6 | 19.3× | beaten on both |
| GLM-4.7 Flash | 74/96 | 77% 68–84 | same | $16.2 | 23.0× | unfit to deploy |
| Claude Sonnet 4.6 | 91/96 | 95% 88–98 | same | $18.8 | 26.7× | beaten on both |
| GLM-5.3 | 85/96 | 89% 81–93 | 93/96 +8 | $22.7 | 32.2× | beaten on both |
| Grok 4.5 | 95/96 | 99% 94–100 | same | $28.3 | 40.1× | beaten on both |
| Qwen3.5 Flash | 53/96 | 55% 45–65 | same | $30.2 | 42.9× | unfit to deploy |
| GLM-5.2 | 91/96 | 95% 88–98 | same | $31.2 | 44.4× | beaten on both |
| Claude Opus 4.6 | 92/96 | 96% 90–98 | 95/96 +3 | $33.6 | 47.7× | beaten on both |
| Kimi K2.5 | 93/96 | 97% 91–99 | same | $35.2 | 50.0× | beaten on both |
| GPT-5.6 Terra Pro | 93/96 | 97% 91–99 | same | $41.5 | 59.0× | beaten on both |
| Grok 4.6 | 94/96 | 98% 93–99 | same | $43.9 | 62.4× | beaten on both |
| Qwen3.8 27B | 90/96 | 94% 87–97 | same | $47.0 | 66.7× | beaten on both |
| DeepSeek V4 Pro | 85/96 | 89% 81–93 | same | $49.2 | 69.9× | beaten on both |
| Kimi K3 | 94/96 | 98% 93–99 | 96/96 +2 | $63.4 | 90.1× | beaten on both |
| Gemini 3.6 Flash | 89/96 | 93% 86–96 | same | $65.4 | 92.9× | beaten on both |
| Qwen3.8 2.4T | 92/96 | 96% 90–98 | same | $79.3 | 112.6× | beaten on both |
| GLM-5 | 79/96 | 82% 73–89 | same | $95.4 | 135.6× | beaten on both |
| Qwen3.8 Max | 90/96 | 94% 87–97 | same | $95.9 | 136.2× | beaten on both |
| Seed 2.1 Turbo | 85/96 | 89% 81–93 | same | $119 | 168.9× | beaten on both |
| Seed 2.0 Code | 85/96 | 89% 81–93 | 86/96 +1 | $146 | 206.8× | beaten on both |
| Claude Opus 5 Fast | 77/96 | 80% 71–87 | 95/96 +18 | $149 | 212.0× | beaten on both |
| GPT-5.6 Sol Pro | 93/96 | 97% 91–99 | same | $192 | 272.8× | beaten on both |
| Muse Spark 1.2 | 0/96 | 0% 0–4 | same | — | — | not measured |