AI Octane Index · BEP Research
Model prices are quoted per token, but you pay per job that comes back right — and a model that needs three tries costs three times. The same basket of work, run through the whole board, spans 267× in what it costs to get finished.
per 10,000 tasks you actually get back correct, retries included. Answered 90 of 96 correctly (87–97% at 95% confidence).
the price of the pick, for 6 more points of observed success (96/96). At this sample size the two are not statistically separable — the intervals overlap, so this premium buys no accuracy we can demonstrate.
Something on this board is cheaper and more accurate than each of them. The worst, GPT-5.6 Sol Pro, costs 267× the pick.
Tap or hover any dot to see which model it is.
Each dot is one model. Left is less accurate, up is more expensive per completed task, so the best place to be is bottom right. Anything inside the shaded region is beaten on both counts at once by something else on this board.
| Industry | Cheapest deployable | Cost per 10,000 completed | Correct |
|---|---|---|---|
| Banking & Finance | DeepSeek V3.2 | $0.32 | 18/18 |
| Customer Service | DeepSeek V4 Flash | $0.31 | 15/15 |
| Legal | GPT-5.6 Luna | $0.73 | 14/15 |
| Medical | DeepSeek V4 Flash | $0.29 | 27/27 |
| Software Engineering | DeepSeek V3.2 | $0.96 | 21/21 |
In 5 of 5 industries the cheapest deployable model is not the board's overall pick. Each industry is a slice of the same basket, so these rest on 15–27 attempts against the 96 behind the board above — a steer, not a ranking.
Calculate the interest accrued. LOAN TERMS: - Principal: $1,000,000 - Rate: 6.00% per annum - Day-count convention: Actual/360 - Accrual period: 45 days Respond with the interest in dollars, rounded to two decimal places, as a number only — no commas, no currency symbol.
7500.00Graded by exact comparison against that answer, so a right number wrapped in an explanation is scored wrong — the prompt asked for a number. Every model on the board gets this same prompt at temperature 0. See all 32 tasks.
Every prompt states an output format — “as a number only” — and the grader compares exactly, so a right answer wrapped in an explanation scores zero. That is a real failure if code reads the output, and no failure at all if a person does. 6 of 27 models are affected; the worst, Llama 4 Maverick, would score 78/96 instead of 54/96 — 24 of its answers were right and rejected on presentation. The ranking above uses the strict score; decide which column is yours.
| Model | Correct | Observed success | Right answer, any format | Cost per 10,000 completed | vs best | Verdict |
|---|---|---|---|---|---|---|
| gpt-oss 120B | 90/96 | 94% 87–97 | same | $0.67 | 1.0× | on the frontier |
| DeepSeek V3.2 | 84/96 | 88% 79–93 | same | $0.69 | 1.0× | beaten on both |
| GPT-5.6 Luna | 95/96 | 99% 94–100 | same | $0.77 | 1.1× | on the frontier |
| DeepSeek V4 Flash | 86/96 | 90% 82–94 | same | $2.74 | 4.1× | beaten on both |
| Gemini 2.5 Flash | 83/96 | 86% 78–92 | same | $2.80 | 4.2× | beaten on both |
| Llama 4 Maverick | 54/96 | 56% 46–66 | 78/96 +24 | $2.85 | 4.2× | format, not accuracy |
| GPT-5.4 mini | 85/96 | 89% 81–93 | same | $3.24 | 4.8× | beaten on both |
| GPT-5.6 Luna Pro | 96/96 | 100% 96–100 | same | $4.71 | 7.0× | on the frontier |
| GPT-5.6 Terra | 93/96 | 97% 91–99 | same | $6.73 | 10.0× | beaten on both |
| Claude Haiku 4.5 | 75/96 | 78% 69–85 | 81/96 +6 | $8.11 | 12.1× | format, not accuracy |
| GLM-5.2 | 92/96 | 96% 90–98 | same | $13.5 | 20.1× | beaten on both |
| GLM-4.7 Flash | 81/96 | 84% 76–90 | same | $16.0 | 23.8× | beaten on both |
| Claude Sonnet 4.6 | 91/96 | 95% 88–98 | 93/96 +2 | $19.5 | 29.1× | beaten on both |
| Grok 4.5 | 93/96 | 97% 91–99 | same | $30.1 | 44.8× | beaten on both |
| Claude Opus 4.6 | 92/96 | 96% 90–98 | 95/96 +3 | $33.0 | 49.1× | beaten on both |
| Qwen3.5 Flash | 54/96 | 56% 46–66 | same | $34.7 | 51.6× | unfit to deploy |
| GPT-5.6 Sol | 93/96 | 97% 91–99 | same | $36.3 | 54.0× | beaten on both |
| Kimi K2.5 | 90/96 | 94% 87–97 | same | $37.1 | 55.2× | beaten on both |
| GPT-5.6 Terra Pro | 93/96 | 97% 91–99 | same | $40.9 | 60.8× | beaten on both |
| Grok 4.6 | 93/96 | 97% 91–99 | same | $42.4 | 63.0× | beaten on both |
| Kimi K3 | 94/96 | 98% 93–99 | 96/96 +2 | $60.4 | 89.9× | beaten on both |
| Gemini 3.6 Flash | 90/96 | 94% 87–97 | same | $66.0 | 98.3× | beaten on both |
| Qwen3.8 2.4T | 89/96 | 93% 86–96 | same | $71.4 | 106.2× | beaten on both |
| Qwen3.8 Max | 91/96 | 95% 88–98 | same | $81.8 | 121.7× | beaten on both |
| GLM-5 | 79/96 | 82% 73–89 | same | $124 | 184.1× | beaten on both |
| Claude Opus 5 Fast | 78/96 | 81% 72–88 | 94/96 +16 | $139 | 206.8× | beaten on both |
| GPT-5.6 Sol Pro | 94/96 | 98% 93–99 | same | $179 | 266.6× | beaten on both |
| DeepSeek V4 Pro | 0/96 | 0% 0–4 | same | — | — | not measured |