AI Octane Index · BEP Research

What a finished task actually costs.

Model prices are quoted per token, but you pay per job that comes back right — and a model that needs three tries costs three times. The same basket of work, run through the whole board, spans 267× in what it costs to get finished.

27 models2,592 graded tests32 tasks each13 August 2026 measured
1600×900, sized for a post

Use this

$0.67
gpt-oss 120B

per 10,000 tasks you actually get back correct, retries included. Answered 90 of 96 correctly (87–97% at 95% confidence).

Most accurate

7.0×
GPT-5.6 Luna Pro

the price of the pick, for 6 more points of observed success (96/96). At this sample size the two are not statistically separable — the intervals overlap, so this premium buys no accuracy we can demonstrate.

What to avoid

24
models beaten on both axes

Something on this board is cheaper and more accurate than each of them. The worst, GPT-5.6 Sol Pro, costs 267× the pick.

$100.00$10.00$1.0060%70%80%90%100%BEATEN ON BOTH24 of 27 models — something here is cheaper and more accurateGPT-5.6 Sol Pro · $179 per 10,000 completed · 94/96 correct · beaten on bothClaude Opus 5 Fast · $139 per 10,000 completed · 78/96 correct · beaten on bothGLM-5 · $124 per 10,000 completed · 79/96 correct · beaten on bothQwen3.8 Max · $81.8 per 10,000 completed · 91/96 correct · beaten on bothQwen3.8 2.4T · $71.4 per 10,000 completed · 89/96 correct · beaten on bothGemini 3.6 Flash · $66.0 per 10,000 completed · 90/96 correct · beaten on bothKimi K3 · $60.4 per 10,000 completed · 94/96 correct · beaten on bothGrok 4.6 · $42.4 per 10,000 completed · 93/96 correct · beaten on bothGPT-5.6 Terra Pro · $40.9 per 10,000 completed · 93/96 correct · beaten on bothKimi K2.5 · $37.1 per 10,000 completed · 90/96 correct · beaten on bothGPT-5.6 Sol · $36.3 per 10,000 completed · 93/96 correct · beaten on bothQwen3.5 Flash · $34.7 per 10,000 completed · 54/96 correct · unfit to deployClaude Opus 4.6 · $33.0 per 10,000 completed · 92/96 correct · beaten on bothGrok 4.5 · $30.1 per 10,000 completed · 93/96 correct · beaten on bothClaude Sonnet 4.6 · $19.5 per 10,000 completed · 91/96 correct · beaten on bothGLM-4.7 Flash · $16.0 per 10,000 completed · 81/96 correct · beaten on bothGLM-5.2 · $13.5 per 10,000 completed · 92/96 correct · beaten on bothClaude Haiku 4.5 · $8.11 per 10,000 completed · 75/96 correct · unfit to deployGPT-5.6 Terra · $6.73 per 10,000 completed · 93/96 correct · beaten on bothGPT-5.6 Luna Pro · $4.71 per 10,000 completed · 96/96 correct · on the frontierGPT-5.6 Luna ProGPT-5.4 mini · $3.24 per 10,000 completed · 85/96 correct · beaten on bothLlama 4 Maverick · $2.85 per 10,000 completed · 54/96 correct · unfit to deployGemini 2.5 Flash · $2.80 per 10,000 completed · 83/96 correct · beaten on bothDeepSeek V4 Flash · $2.74 per 10,000 completed · 86/96 correct · beaten on bothGPT-5.6 Luna · $0.77 per 10,000 completed · 95/96 correct · on the frontierGPT-5.6 LunaDeepSeek V3.2 · $0.69 per 10,000 completed · 84/96 correct · beaten on bothgpt-oss 120B · $0.67 per 10,000 completed · 90/96 correct · on the frontiergpt-oss 120Bshare of tasks answered correctly → bettercost per 10,000 completed tasks ↓ better

Tap or hover any dot to see which model it is.

Each dot is one model. Left is less accurate, up is more expensive per completed task, so the best place to be is bottom right. Anything inside the shaded region is beaten on both counts at once by something else on this board.

Cheapest model that can do the work, by industry

IndustryCheapest deployableCost per 10,000 completedCorrect
Banking & FinanceDeepSeek V3.2$0.3218/18
Customer ServiceDeepSeek V4 Flash$0.3115/15
LegalGPT-5.6 Luna$0.7314/15
MedicalDeepSeek V4 Flash$0.2927/27
Software EngineeringDeepSeek V3.2$0.9621/21

In 5 of 5 industries the cheapest deployable model is not the board's overall pick. Each industry is a slice of the same basket, so these rest on 15–27 attempts against the 96 behind the board above — a steer, not a ranking.

What one of the 32 tasks looks like

Accrue interest under a stated day-count convention Banking & Finance
Calculate the interest accrued.

LOAN TERMS:
  - Principal: $1,000,000
  - Rate: 6.00% per annum
  - Day-count convention: Actual/360
  - Accrual period: 45 days

Respond with the interest in dollars, rounded to two decimal places, as a number
only — no commas, no currency symbol.
Accepted answer7500.00

Graded by exact comparison against that answer, so a right number wrapped in an explanation is scored wrong — the prompt asked for a number. Every model on the board gets this same prompt at temperature 0. See all 32 tasks.

Every model measured

Every prompt states an output format — “as a number only” — and the grader compares exactly, so a right answer wrapped in an explanation scores zero. That is a real failure if code reads the output, and no failure at all if a person does. 6 of 27 models are affected; the worst, Llama 4 Maverick, would score 78/96 instead of 54/96 — 24 of its answers were right and rejected on presentation. The ranking above uses the strict score; decide which column is yours.

ModelCorrectObserved successRight answer, any formatCost per 10,000 completedvs bestVerdict
gpt-oss 120B90/9694% 87–97same$0.671.0×on the frontier
DeepSeek V3.284/9688% 79–93same$0.691.0×beaten on both
GPT-5.6 Luna95/9699% 94–100same$0.771.1×on the frontier
DeepSeek V4 Flash86/9690% 82–94same$2.744.1×beaten on both
Gemini 2.5 Flash83/9686% 78–92same$2.804.2×beaten on both
Llama 4 Maverick54/9656% 46–6678/96 +24$2.854.2×format, not accuracy
GPT-5.4 mini85/9689% 81–93same$3.244.8×beaten on both
GPT-5.6 Luna Pro96/96100% 96–100same$4.717.0×on the frontier
GPT-5.6 Terra93/9697% 91–99same$6.7310.0×beaten on both
Claude Haiku 4.575/9678% 69–8581/96 +6$8.1112.1×format, not accuracy
GLM-5.292/9696% 90–98same$13.520.1×beaten on both
GLM-4.7 Flash81/9684% 76–90same$16.023.8×beaten on both
Claude Sonnet 4.691/9695% 88–9893/96 +2$19.529.1×beaten on both
Grok 4.593/9697% 91–99same$30.144.8×beaten on both
Claude Opus 4.692/9696% 90–9895/96 +3$33.049.1×beaten on both
Qwen3.5 Flash54/9656% 46–66same$34.751.6×unfit to deploy
GPT-5.6 Sol93/9697% 91–99same$36.354.0×beaten on both
Kimi K2.590/9694% 87–97same$37.155.2×beaten on both
GPT-5.6 Terra Pro93/9697% 91–99same$40.960.8×beaten on both
Grok 4.693/9697% 91–99same$42.463.0×beaten on both
Kimi K394/9698% 93–9996/96 +2$60.489.9×beaten on both
Gemini 3.6 Flash90/9694% 87–97same$66.098.3×beaten on both
Qwen3.8 2.4T89/9693% 86–96same$71.4106.2×beaten on both
Qwen3.8 Max91/9695% 88–98same$81.8121.7×beaten on both
GLM-579/9682% 73–89same$124184.1×beaten on both
Claude Opus 5 Fast78/9681% 72–8894/96 +16$139206.8×beaten on both
GPT-5.6 Sol Pro94/9698% 93–99same$179266.6×beaten on both
DeepSeek V4 Pro0/960% 0–4samenot measured
AI OCTANE INDEX · BEP RESEARCHWhat a finished task actually costs.27 models · 32 identical tasks · priced on what it takes to get a right answer outNOT BEATEN ON PRICE OR ACCURACYgpt-oss 120B$0.6790/96 correctGPT-5.6 Luna$0.7795/96 correctGPT-5.6 Luna Pro$4.7196/96 correctMOST EXPENSIVEGPT-5.6 Sol Pro$17994/96 correct2 cheaper models answered more of the basket correctly.267×spread in cost per completed task3 of 27 models are not beaten on both price and accuracy · temperature 0 · 13 August 2026tools.bepresearch.com/octane