AI Octane Index · BEP Research

What a finished task actually costs.

Model prices are quoted per token, but you pay per job that comes back right — and a model that needs three tries costs three times. The same basket of work, run through the whole board, spans 273× in what it costs to get finished.

33 models3,168 graded tests32 tasks each22 August 2026 measured
1600×900, sized for a post

Lowest measured cost

$0.70
gpt-oss 120B

weighted cost per 10,000 successful attempts in this benchmark, including the cost of failed attempts. Answered 92 of 96 correctly (90–98% at 95% confidence).

Highest observed accuracy

13.4×
GPT-5.6 Luna Pro

the price of the pick, for 4 more points of observed success (96/96). This small sample does not establish the size of an accuracy premium. Overlapping intervals alone are not a significance test; a larger paired evaluation is needed.

Dominated in this sample

30
models beaten on both axes

Something on this board is cheaper and more accurate than each of them. The worst, GPT-5.6 Sol Pro, costs 273× the pick.

$100.00$10.00$1.0050%60%70%80%90%100%BEATEN ON BOTH30 of 33 models — something here is cheaper and more accurateGPT-5.6 Sol Pro · $192 per 10,000 completed · 93/96 correct · beaten on bothClaude Opus 5 Fast · $149 per 10,000 completed · 77/96 correct · beaten on bothSeed 2.0 Code · $146 per 10,000 completed · 85/96 correct · beaten on bothSeed 2.1 Turbo · $119 per 10,000 completed · 85/96 correct · beaten on bothQwen3.8 Max · $95.9 per 10,000 completed · 90/96 correct · beaten on bothGLM-5 · $95.4 per 10,000 completed · 79/96 correct · beaten on bothQwen3.8 2.4T · $79.3 per 10,000 completed · 92/96 correct · beaten on bothGemini 3.6 Flash · $65.4 per 10,000 completed · 89/96 correct · beaten on bothKimi K3 · $63.4 per 10,000 completed · 94/96 correct · beaten on bothDeepSeek V4 Pro · $49.2 per 10,000 completed · 85/96 correct · beaten on bothQwen3.8 27B · $47.0 per 10,000 completed · 90/96 correct · beaten on bothGrok 4.6 · $43.9 per 10,000 completed · 94/96 correct · beaten on bothGPT-5.6 Terra Pro · $41.5 per 10,000 completed · 93/96 correct · beaten on bothKimi K2.5 · $35.2 per 10,000 completed · 93/96 correct · beaten on bothClaude Opus 4.6 · $33.6 per 10,000 completed · 92/96 correct · beaten on bothGLM-5.2 · $31.2 per 10,000 completed · 91/96 correct · beaten on bothQwen3.5 Flash · $30.2 per 10,000 completed · 53/96 correct · unfit to deployGrok 4.5 · $28.3 per 10,000 completed · 95/96 correct · beaten on bothGLM-5.3 · $22.7 per 10,000 completed · 85/96 correct · beaten on bothClaude Sonnet 4.6 · $18.8 per 10,000 completed · 91/96 correct · beaten on bothGLM-4.7 Flash · $16.2 per 10,000 completed · 74/96 correct · unfit to deployGPT-5.6 Terra · $13.6 per 10,000 completed · 93/96 correct · beaten on bothGPT-5.6 Sol · $11.0 per 10,000 completed · 93/96 correct · beaten on bothGemini 3.7 Flash · $9.88 per 10,000 completed · 92/96 correct · beaten on bothGPT-5.6 Luna Pro · $9.44 per 10,000 completed · 96/96 correct · on the frontierGPT-5.6 Luna ProDeepSeek V4 Flash · $8.51 per 10,000 completed · 85/96 correct · beaten on bothClaude Haiku 4.5 · $8.25 per 10,000 completed · 75/96 correct · unfit to deployGPT-5.4 mini · $4.21 per 10,000 completed · 79/96 correct · beaten on bothLlama 4 Maverick · $3.36 per 10,000 completed · 56/96 correct · unfit to deployGemini 2.5 Flash · $3.22 per 10,000 completed · 84/96 correct · beaten on bothGPT-5.6 Luna · $1.64 per 10,000 completed · 94/96 correct · on the frontierGPT-5.6 LunaDeepSeek V3.2 · $0.73 per 10,000 completed · 83/96 correct · beaten on bothgpt-oss 120B · $0.70 per 10,000 completed · 92/96 correct · on the frontiergpt-oss 120Bshare of tasks answered correctly → bettercost per 10,000 completed tasks ↓ better

Tap or hover any dot to see which model it is.

Each dot is one model. Left is less accurate, up is more expensive per completed task, so the best place to be is bottom right. Anything inside the shaded region is beaten on both counts at once by something else on this board.

Cheapest model that can do the work, by industry

IndustryLowest cost above the test thresholdCost per 10,000 completedCorrect
Banking & Financegpt-oss 120B$0.3818/18
Customer ServiceDeepSeek V4 Flash$0.3115/15
Legalgpt-oss 120B$1.1812/15
MedicalDeepSeek V4 Flash$0.2827/27
Software EngineeringDeepSeek V3.2$1.0021/21

In 3 of 5 industries the cheapest deployable model is not the board's overall pick. Each industry is a slice of the same basket, so these rest on 15–27 attempts against the 96 behind the board above — a steer, not a ranking.

What one of the 32 tasks looks like

Accrue interest under a stated day-count convention Banking & Finance
Calculate the interest accrued.

LOAN TERMS:
  - Principal: $1,000,000
  - Rate: 6.00% per annum
  - Day-count convention: Actual/360
  - Accrual period: 45 days

Respond with the interest in dollars, rounded to two decimal places, as a number
only — no commas, no currency symbol.
Accepted answer7500.00

Graded by exact comparison against that answer, so a right number wrapped in an explanation is scored wrong — the prompt asked for a number. Every model on the board gets this same prompt at temperature 0. See all 32 tasks.

Every model measured

Every prompt states an output format — “as a number only” — and the grader compares exactly, so a right answer wrapped in an explanation scores zero. That is a real failure if code reads the output, and no failure at all if a person does. 7 of 33 models are affected; the worst, Llama 4 Maverick, would score 81/96 instead of 56/96 — 25 of its answers were right and rejected on presentation. The ranking above uses the strict score; decide which column is yours.

ModelCorrectObserved successRight answer, any formatCost per 10,000 completedvs bestVerdict
gpt-oss 120B92/9696% 90–98same$0.701.0×on the frontier
DeepSeek V3.283/9686% 78–92same$0.731.0×beaten on both
GPT-5.6 Luna94/9698% 93–99same$1.642.3×on the frontier
Gemini 2.5 Flash84/9688% 79–93same$3.224.6×beaten on both
Llama 4 Maverick56/9658% 48–6881/96 +25$3.364.8×format, not accuracy
GPT-5.4 mini79/9682% 73–89same$4.216.0×beaten on both
Claude Haiku 4.575/9678% 69–8581/96 +6$8.2511.7×format, not accuracy
DeepSeek V4 Flash85/9689% 81–93same$8.5112.1×beaten on both
GPT-5.6 Luna Pro96/96100% 96–100same$9.4413.4×on the frontier
Gemini 3.7 Flash92/9696% 90–98same$9.8814.0×beaten on both
GPT-5.6 Sol93/9697% 91–99same$11.015.6×beaten on both
GPT-5.6 Terra93/9697% 91–99same$13.619.3×beaten on both
GLM-4.7 Flash74/9677% 68–84same$16.223.0×unfit to deploy
Claude Sonnet 4.691/9695% 88–98same$18.826.7×beaten on both
GLM-5.385/9689% 81–9393/96 +8$22.732.2×beaten on both
Grok 4.595/9699% 94–100same$28.340.1×beaten on both
Qwen3.5 Flash53/9655% 45–65same$30.242.9×unfit to deploy
GLM-5.291/9695% 88–98same$31.244.4×beaten on both
Claude Opus 4.692/9696% 90–9895/96 +3$33.647.7×beaten on both
Kimi K2.593/9697% 91–99same$35.250.0×beaten on both
GPT-5.6 Terra Pro93/9697% 91–99same$41.559.0×beaten on both
Grok 4.694/9698% 93–99same$43.962.4×beaten on both
Qwen3.8 27B90/9694% 87–97same$47.066.7×beaten on both
DeepSeek V4 Pro85/9689% 81–93same$49.269.9×beaten on both
Kimi K394/9698% 93–9996/96 +2$63.490.1×beaten on both
Gemini 3.6 Flash89/9693% 86–96same$65.492.9×beaten on both
Qwen3.8 2.4T92/9696% 90–98same$79.3112.6×beaten on both
GLM-579/9682% 73–89same$95.4135.6×beaten on both
Qwen3.8 Max90/9694% 87–97same$95.9136.2×beaten on both
Seed 2.1 Turbo85/9689% 81–93same$119168.9×beaten on both
Seed 2.0 Code85/9689% 81–9386/96 +1$146206.8×beaten on both
Claude Opus 5 Fast77/9680% 71–8795/96 +18$149212.0×beaten on both
GPT-5.6 Sol Pro93/9697% 91–99same$192272.8×beaten on both
Muse Spark 1.20/960% 0–4samenot measured
AI OCTANE INDEX · BEP RESEARCHWhat a finished task actually costs.33 models · 32 identical tasks · priced on what it takes to get a right answer outNOT BEATEN ON PRICE OR ACCURACYgpt-oss 120B$0.7092/96 correctGPT-5.6 Luna$1.6494/96 correctGPT-5.6 Luna Pro$9.4496/96 correctMOST EXPENSIVEGPT-5.6 Sol Pro$19293/96 correct5 cheaper models answered more of the basket correctly.273×spread in cost per completed task3 of 33 models are not beaten on both price and accuracy · temperature 0 · 22 August 2026tools.bepresearch.com/octane