BEP Research Research preview

34 models, priced per finished task.

The current board — every model run on the same 32 tasks, graded the same way, and costed on what it actually spent to get work accepted. The 10 models at the top are doing work we cannot tell apart, at a 117× spread in price.

Price is what you pay for tokens. Cost is what you pay for mistakes.

benchmark 0.3.0 · pricing 2026-07-09 · billed cost · baseline gpt-oss 120B · 32 tasks · computed 2026-08-22 15:48
34 AI models tested3,264 graded testsAugust 2026 last updated
What are you trying to automate?
How much of it, and what does a mistake cost?
Tasks per month
A wrong answer that reaches a customer costs us
Most teams underestimate this. $200 is a reasonable starting point for a customer-facing mistake that needs finding, apologising for and redoing.
Accuracy we require

Workload calculator

The benchmark measures what an attempt costs and how often it is accepted. Everything else that decides your bill is yours: volume, review time, what a wrong answer costs when it reaches a customer. Supply those and the ranking can change — which is the point.

Retries assume a failed task fails again, because that is what we measured — asking a second time usually returns the same wrong answer at twice the price. Why that is, and when it is not true →

Your workload assumption
Human review assumption
Risk and retries assumption

Success rates and cost per attempt are measured on the benchmark tasks shown below. Volume, review time, error cost, detection rate and retry behaviour are assumptions you supplied. Totals combine both and are estimates, not quotes.

The evidence

Everything below is how that recommendation was reached: what each model cost, how often a grader accepted its output, and where the numbers are too thin to act on. The Octane score explains the ranking; it is not the product.

$0.70 per 10,000 completed tasks · gpt-oss 120B · 272.8× cheaper than GPT-5.6 Sol Pro
baseline
observed success of that option
96%
417 of 10,000 still need redoing
weakest vertical
80%
Legal — weakest measured vertical
widest gap in one vertical
480×
how much the choice is worth
kinds of work tested
5
3,264 graded tests

Which model won each kind of work

These cards reflect performance on the specific benchmark tasks listed under “What we test” — they are starting points for a model evaluation, not blanket recommendations for an entire industry. The badge is the observed success rate on those tasks, with the run counts beneath it; not a claim about production reliability. Each card leads with the highest-scoring model above the 80% threshold, then names a lower-cost option that also cleared it. Where no model reached the threshold, the card says so instead of recommending one.

Banking & Finance
100%observed success
18 / 18 benchmark runs
gpt-oss 120B
$0.38 per 10,000 successful tasks
Lowest measured cost above the threshold
Customer Service
100%observed success
15 / 15 benchmark runs
DeepSeek V4 Flash
$0.31 per 10,000 successful tasks
Lowest measured cost above the threshold
Legal
100%observed success
15 / 15 benchmark runs
GPT-5.6 Luna Pro
$9.13 per 10,000 successful tasks
Lower-cost option: gpt-oss 120B — 12/15 (80%), $1.18, 7.8× cheaper
Medical
100%observed success
27 / 27 benchmark runs
DeepSeek V4 Flash
$0.28 per 10,000 successful tasks
Lowest measured cost above the threshold
Software Engineering
100%observed success
21 / 21 benchmark runs
DeepSeek V3.2
$1.00 per 10,000 successful tasks
Lowest measured cost above the threshold

Sticker price is not the price

A pricing page sells you a rate per million tokens. That rate is one of three terms, and it is the only one published. A model can advertise the lowest rate, write two or three times as many tokens to answer the same question, fail a share of the time — and finish as the most expensive option here.

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

what the price page implies →$ per 10,000 completed tasksgpt-oss 120B$0.00$0.59$0.61$0.70DeepSeek V3.2$0.00$0.52$0.60$0.73GPT-5.6 Luna$0.00$1.56$1.60$1.64Gemini 2.5 Flash$0.00$2.72$3.11$3.22Llama 4 Maverick$0.00$1.42$2.44$3.36GPT-5.4 mini$0.00$2.60$3.16$4.21Claude Haiku 4.5$0.00$4.94$6.33$8.25DeepSeek V4 Flash$0.00$6.02$6.80$8.51GPT-5.6 Luna Pro$0.00$9.25$9.25$9.44Gemini 3.7 Flash$0.00$8.88$9.27$9.88GPT-5.6 Sol$0.00$10.20$10.53$10.96GPT-5.6 Terra$0.00$12.78$13.19$13.62GLM-4.7 Flash$0.00$8.96$11.63$16.20Claude Sonnet 4.6$0.00$16.42$17.33$18.79GLM-5.3$0.00$18.19$20.55$22.68Grok 4.5$0.00$26.28$26.55$28.25Qwen3.5 Flash$0.00$5.59$10.13$30.22GLM-5.2$0.00$26.24$27.68$31.22Claude Opus 4.6$0.00$31.47$32.84$33.57Kimi K2.5$0.00$30.56$31.54$35.22GPT-5.6 Terra Pro$0.00$39.03$40.29$41.54Grok 4.6$0.00$38.67$39.49$43.90Qwen3.8 27B$0.00$39.31$41.93$46.96DeepSeek V4 Pro$0.00$36.74$41.49$49.23Kimi K3$0.00$57.56$58.79$63.40Gemini 3.6 Flash$0.00$55.90$60.30$65.39Qwen3.8 2.4T$0.00$61.91$64.60$79.28GLM-5$0.00$48.02$58.36$95.43Qwen3.8 Max$0.00$69.45$74.08$95.86Seed 2.1 Turbo$0.00$74.69$84.35$118.92Seed 2.0 Code$0.00$98.01$110.69$145.61Claude Opus 5 Fast$0.00$108.40$135.15$149.22GPT-5.6 Sol Pro$0.00$177.81$183.54$192.06$0$50$101$151$202

Each bar adds one thing the price page does not tell you. The faintest is the advertised rate on this exact workload, priced as if the model were as concise as the leanest one measured. Then the tokens it really wrote. Then the attempts you paid for and could not use. Then your actual work mix, since the index weights every vertical equally instead of letting a model's weak vertical be diluted by its strong one. Only the faintest bar is visible when you pick a model.

Overall

The leading model in each panel sets 100; every other grade is its share of that leader's successful work per dollar. Log axis — a real panel spans orders of magnitude. Whiskers are 95% clustered bootstrap intervals, resampled by vertical, then task, then repetition. Raw baseline-relative scores are in the CSV and JSON exports.

gpt-oss 120BClaude Haiku 4.5Claude Sonnet 4.6Claude Opus 4.6Qwen3.5 FlashLlama 4 MaverickDeepSeek V3.2Kimi K2.5GLM-5GLM-5.2GLM-4.7 FlashGPT-5.4 miniGPT-5.6 LunaGPT-5.6 Luna ProGPT-5.6 TerraGPT-5.6 SolGemini 2.5 FlashDeepSeek V4 FlashKimi K3Qwen3.8 MaxGrok 4.6Grok 4.5Claude Opus 5 FastGemini 3.6 FlashDeepSeek V4 ProQwen3.8 2.4TGPT-5.6 Terra ProGPT-5.6 Sol ProGLM-5.3Qwen3.8 27BGemini 3.7 FlashSeed 2.1 TurboSeed 2.0 CodeMuse Spark 1.2
0.010.101.0010.01001,000100 = scope leadergpt-oss 120B100DeepSeek V3.296.1GPT-5.6 Luna43.1Gemini 2.5 Flash21.9Llama 4 Maverick21.056/96 — below thresholdGPT-5.4 mini16.7Claude Haiku 4.58.5475/96 — below thresholdDeepSeek V4 Flash8.27GPT-5.6 Luna Pro7.45Gemini 3.7 Flash7.13GPT-5.6 Sol6.43GPT-5.6 Terra5.17GLM-4.7 Flash4.3574/96 — below thresholdClaude Sonnet 4.63.75GLM-5.33.10Grok 4.52.49Qwen3.5 Flash2.3353/96 — below thresholdGLM-5.22.25Claude Opus 4.62.10Kimi K2.52.00GPT-5.6 Terra Pro1.69Grok 4.61.60Qwen3.8 27B1.50DeepSeek V4 Pro1.43Kimi K31.11Gemini 3.6 Flash1.08Qwen3.8 2.4T0.89GLM-50.74Qwen3.8 Max0.73Seed 2.1 Turbo0.59Seed 2.0 Code0.48Claude Opus 5 Fast0.47GPT-5.6 Sol Pro0.37Muse Spark 1.2no successes — score unavailable
Why the overall leader differs from the vertical leaders

gpt-oss 120B ranks first overall because its lead in Banking & Finance, Legal outweighs DeepSeek V4 Flash's lead in Customer Service, Medical under the benchmark's equal-weight methodology — each vertical counts the same regardless of how many tasks or attempts it holds. Weight the verticals to your own workload in the calculator above and the ranking can change.

Observations
  • The choice of model matters most in Software Engineering, where measured cost per successful task spans 480x.
  • GPT-5.6 Luna Pro completed every benchmark run (96/96).

By vertical

Every panel shares one axis, so they are directly comparable. This is the cut that matters — cost-effectiveness is a property of a model on a kind of work, not of a model alone.

A hollow, struck-through dot marks a model that recorded below 80% observed success on that vertical. It is excluded from the recommendation, not judged unusable — cost per successful task already prices retries, but it assumes a failure is free to detect and that retrying eventually works. Neither held in these measurements: most model-task pairs returned byte-identical text on every repetition, so a failed attempt repeated rather than varied.

Banking & Finance
0.010.101.0010.01001,000gpt-oss 120B100DeepSeek V3.297.4DeepSeek V4 Flash88.2GPT-5.6 Luna52.4Gemini 2.5 Flash46.5GPT-5.4 mini28.9Claude Haiku 4.519.7DeepSeek V4 Pro17.4GLM-4.7 Flash12.9GPT-5.6 Sol11.6Qwen3.5 Flash9.31Llama 4 Maverick8.81GPT-5.6 Terra7.56GPT-5.6 Luna Pro6.31Gemini 3.7 Flash5.78Qwen3.8 27B5.54Claude Sonnet 4.64.48GLM-5.34.45GLM-5.24.07Kimi K2.53.19GLM-52.95Grok 4.52.59Qwen3.8 2.4T2.56Qwen3.8 Max2.51Seed 2.1 Turbo1.96Grok 4.61.73GPT-5.6 Terra Pro1.68Claude Opus 4.61.26Gemini 3.6 Flash0.89Kimi K30.77Seed 2.0 Code0.74Claude Opus 5 Fast0.60GPT-5.6 Sol Pro0.52Muse Spark 1.2no successes — score unavailable
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 18/18 successful · 100% observed.
  • Measured cost per successful task spans 194x across the field, from gpt-oss 120B to GPT-5.6 Sol Pro.
Customer Service
0.010.101.0010.01001,000DeepSeek V4 Flash100DeepSeek V3.286.6gpt-oss 120B60.1Gemini 2.5 Flash50.8GPT-5.6 Luna48.2GPT-5.4 mini17.4DeepSeek V4 Pro16.5GLM-4.7 Flash15.0Llama 4 Maverick12.5GPT-5.6 Sol6.21Gemini 3.7 Flash5.97GPT-5.6 Luna Pro5.71GPT-5.6 Terra5.49Qwen3.8 27B5.31GLM-5.24.84Qwen3.5 Flash4.82GLM-53.28Kimi K2.52.65GLM-5.32.61Qwen3.8 2.4T2.29Qwen3.8 Max2.21Claude Haiku 4.52.07Claude Sonnet 4.62.04Grok 4.51.84GPT-5.6 Terra Pro1.46Kimi K31.31Grok 4.61.20Seed 2.1 Turbo1.19Claude Opus 4.61.04Gemini 3.6 Flash0.96Seed 2.0 Code0.87GPT-5.6 Sol Pro0.31Claude Opus 5 Fast0.28Muse Spark 1.2no successes — score unavailable
  • DeepSeek V4 Flash recorded the lowest measured cost among models above the threshold, at 15/15 successful · 100% observed.
  • Measured cost per successful task spans 362x across the field, from DeepSeek V4 Flash to Claude Opus 5 Fast.
Legal
0.010.101.0010.01001,000gpt-oss 120B100DeepSeek V3.277.9GPT-5.6 Luna77.9DeepSeek V4 Flash76.1Gemini 2.5 Flash41.1Llama 4 Maverick24.0Claude Haiku 4.515.2GPT-5.6 Luna Pro12.9GPT-5.4 mini12.6GPT-5.6 Sol10.5GPT-5.6 Terra9.61Gemini 3.7 Flash9.28GLM-4.7 Flash8.06Qwen3.5 Flash7.59GLM-5.26.19DeepSeek V4 Pro5.36Claude Sonnet 4.64.58Qwen3.8 27B4.44Claude Opus 4.64.21Grok 4.53.39GLM-5.33.34GPT-5.6 Terra Pro2.82Kimi K2.52.38Grok 4.61.71Qwen3.8 2.4T1.70Gemini 3.6 Flash1.47Kimi K31.33GLM-50.84Qwen3.8 Max0.76GPT-5.6 Sol Pro0.57Seed 2.0 Code0.55Claude Opus 5 Fast0.51Seed 2.1 Turbo0.46Muse Spark 1.2no successes — score unavailable
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 12/15 successful · 80% observed.
  • Measured cost per successful task spans 216x across the field, from gpt-oss 120B to Seed 2.1 Turbo.
Medical
0.010.101.0010.01001,000DeepSeek V4 Flash100gpt-oss 120B92.6DeepSeek V3.268.5GPT-5.6 Luna36.3Gemini 2.5 Flash20.6GPT-5.4 mini18.4DeepSeek V4 Pro17.9Llama 4 Maverick8.00GLM-4.7 Flash7.03Claude Haiku 4.56.46Qwen3.5 Flash6.28GPT-5.6 Sol5.69Gemini 3.7 Flash5.43Qwen3.8 27B5.37GPT-5.6 Luna Pro4.71GPT-5.6 Terra4.38GLM-5.24.05GLM-5.33.99Qwen3.8 Max2.53Claude Sonnet 4.62.46GLM-52.42Qwen3.8 2.4T2.38Kimi K2.52.09Grok 4.51.60Claude Opus 4.61.41Grok 4.61.37Kimi K31.35Seed 2.1 Turbo1.32GPT-5.6 Terra Pro1.04Seed 2.0 Code0.96Gemini 3.6 Flash0.88GPT-5.6 Sol Pro0.28Claude Opus 5 Fast0.27Muse Spark 1.2no successes — score unavailable
  • DeepSeek V4 Flash recorded the lowest measured cost among models above the threshold, at 27/27 successful · 100% observed.
  • Measured cost per successful task spans 367x across the field, from DeepSeek V4 Flash to Claude Opus 5 Fast.
Software Engineering
0.010.101.0010.01001,000DeepSeek V3.2100gpt-oss 120B86.9Llama 4 Maverick61.8GPT-5.6 Luna22.1GPT-5.4 mini14.0Gemini 2.5 Flash9.57Claude Haiku 4.58.01Gemini 3.7 Flash5.04GPT-5.6 Luna Pro4.81GPT-5.6 Sol3.28Claude Sonnet 4.62.99GPT-5.6 Terra2.57DeepSeek V4 Flash2.50GLM-5.31.97GLM-4.7 Flash1.74Grok 4.51.73Claude Opus 4.61.66Grok 4.61.21Kimi K2.51.11GPT-5.6 Terra Pro1.05GLM-5.20.87Qwen3.5 Flash0.83Kimi K30.74Gemini 3.6 Flash0.71Qwen3.8 27B0.53DeepSeek V4 Pro0.46Claude Opus 5 Fast0.42Seed 2.1 Turbo0.36Qwen3.8 Max0.35Qwen3.8 2.4T0.35GLM-50.33Seed 2.0 Code0.25GPT-5.6 Sol Pro0.21Muse Spark 1.2no successes — score unavailable
  • DeepSeek V3.2 recorded the lowest measured cost among models above the threshold, at 21/21 successful · 100% observed.
  • Measured cost per successful task spans 480x across the field, from DeepSeek V3.2 to GPT-5.6 Sol Pro.

The frontier

Observed success against cost per successful task. Bottom-right is the good corner: high observed success at low cost. Position carries the reading, not colour — every point is directly labelled.

$0.00001$0.00010$0.00100$0.0100$0.10000%25%50%75%100%success rate → bettercost per success ↓ bettergpt-oss 120BDeepSeek V3.2GPT-5.6 LunaGemini 2.5 FlashLlama 4 MaverickGPT-5.4 miniClaude Haiku 4.5DeepSeek V4 FlashGPT-5.6 Luna ProGemini 3.7 FlashGPT-5.6 SolGPT-5.6 TerraGLM-4.7 FlashClaude Sonnet 4.6GLM-5.3Grok 4.5Qwen3.5 FlashGLM-5.2Claude Opus 4.6Kimi K2.5GPT-5.6 Terra ProGrok 4.6Qwen3.8 27BDeepSeek V4 ProKimi K3Gemini 3.6 FlashQwen3.8 2.4TGLM-5Qwen3.8 MaxSeed 2.1 TurboSeed 2.0 CodeClaude Opus 5 FastGPT-5.6 Sol Pro

Methodology

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

Total billed cost includes input tokens, output tokens, cached tokens where applicable, reasoning tokens where separately billed, failed attempts, and retries where benchmarked. Work you paid for and could not use stays in the numerator.

Full methodology, grader definitions, confidence rules and limitations →

What we test

Every task behind these numbers, with how it is graded and what it weighs. Full prompts are deliberately unpublished — printed prompts end up in training data and the benchmark decays — but per-task results are in the CSV and JSON exports, so any figure on this page traces to graded attempts.

Banking & Finance 6 tasks

TaskGraded byWeightVersion
Accrue interest under a stated day-count conventionexact answer2v1
Convert a nominal rate to an effective annual ratestructured output, field-checked1.5v1
Refuse to settle an FX conversion at the wrong date's rateexact answer2v1
Assign a KYC risk tierexact answer1v1
Allocate a partial payment down a contractual waterfallstructured output, field-checked1.5v1
Extract structured fields from a bank statement linestructured output, field-checked1v1

Customer Service 5 tasks

TaskGraded byWeightVersion
Apply a return policy where the exception overrides the deadlineexact answer2v1
Apply a multi-condition escalation ruleexact answer1.5v1
Compute a prorated refund on cancellationstructured output, field-checked1.5v1
Decide refund eligibility against a policystructured output, field-checked1v1
Classify a support ticket into a routing intentexact answer1v1

Legal 5 tasks

TaskGraded byWeightVersion
Resolve conflicting governing-law clauses via a precedence rulestructured output, field-checked1.5v1
Extract governing law and venue from a contract clausestructured output, field-checked1v1
Apply a liability cap with a carve-out and an absolute exclusionexact answer2v1
Compute a notice expiry with exclusion and weekend roll-forwardstructured output, field-checked1.5v1
Earliest termination date under an initial-term barexact answer2v1

Medical 9 tasks

TaskGraded byWeightVersion
Apply a drug interaction only when its stated condition holdsexact answer2v1
Identify the ICD-10 code for a described conditionexact answer1v2
Convert a weight-based infusion order into a pump rateexact answer2v1
Extract a drug interaction into structured formstructured output, field-checked1v1
Apply a maximum-dose ceiling before converting to volumeexact answer2v1
Refuse to fill a weight-based order when the weight is missingexact answer2v1
Convert a weight-based dose into a dispensing volumestructured output, field-checked1.5v1
Interpret a lab value against a sex-specific reference rangeexact answer1.5v1
Adjust a dose for renal function using Cockcroft-Gaultexact answer2v1

Software Engineering 7 tasks

TaskGraded byWeightVersion
Add months to a date, clamping to the end of the monthcode executed against tests2v1
Implement order-preserving deduplicationcode executed against tests1.5v1
Predict the output of an order-preserving deduplicationexact answer1v1
Merge overlapping intervals, including touching endpointscode executed against tests1.5v1
Count divisible pairs at a scale that punishes brute forcecode executed against tests2v1
Parse durations, honouring a buried error-handling contractcode executed against tests1.5v1
Subtract one set of half-open intervals from anothercode executed against tests2v1

Benchmark your own workload

This public basket is 23 short, single-shot tasks. Your work is not these tasks. The same machinery runs against representative samples of your own workload — extraction from your documents, your support tickets, your code — and returns model comparison, total-cost analysis under your review and error costs, observed success by task type, routing recommendations, and the caveats that apply.

Useful for model selection, vendor evaluation, cost optimisation, routing design and procurement support. Get in touch to benchmark your workload →

All figures

Task-level rows are omitted here for length. The CSV and JSON exports carry them in full.

ScopeVerticalModelObserved successCost per successful taskCost 95% CIOctaneOctane 95% CITotal billed cost
overallgpt-oss 120B92/96 · 96%$0.0000704$0.0000319–$0.00012100$0.00565
overallDeepSeek V3.283/96 · 86%$0.000073296.146.3–192$0.00497
overallGPT-5.6 Luna94/96 · 98%$0.0001643.123.2–86.0$0.0150
overallGemini 2.5 Flash84/96 · 88%$0.0003221.910.9–74.4$0.0261
overallLlama 4 Maverick56/96 · 58%$0.0003421.05.42–47.9$0.0137
overallGPT-5.4 mini79/96 · 82%$0.0004216.77.28–29.7$0.0250
overallClaude Haiku 4.575/96 · 78%$0.000828.542.76–17.9$0.0474
overallDeepSeek V4 Flash85/96 · 89%$0.000858.271.63–138$0.0578
overallGPT-5.6 Luna Pro96/96 · 100%$0.000947.454.50–13.0$0.0888
overallGemini 3.7 Flash92/96 · 96%$0.000997.134.84–11.4$0.0853
overallGPT-5.6 Sol93/96 · 97%$0.001106.433.48–12.1$0.0979
overallGPT-5.6 Terra93/96 · 97%$0.001365.172.70–10.2$0.1226
overallGLM-4.7 Flash74/96 · 77%$0.001624.351.36–18.7$0.0860
overallClaude Sonnet 4.691/96 · 95%$0.001883.752.34–5.70$0.1577
overallGLM-5.385/96 · 89%$0.002273.101.54–5.99$0.1746
overallGrok 4.595/96 · 99%$0.002832.491.75–3.76$0.2522
overallQwen3.5 Flash53/96 · 55%$0.003022.331.19–17.4$0.0537
overallGLM-5.291/96 · 95%$0.003122.250.85–7.50$0.2519
overallClaude Opus 4.692/96 · 96%$0.003362.101.21–3.60$0.3021
overallKimi K2.593/96 · 97%$0.003522.001.20–3.85$0.2933
overallGPT-5.6 Terra Pro93/96 · 97%$0.004151.691.03–2.91$0.3747
overallGrok 4.694/96 · 98%$0.004391.601.18–2.27$0.3712
overallQwen3.8 27B90/96 · 94%$0.004701.500.47–7.97$0.3774
overallDeepSeek V4 Pro85/96 · 89%$0.004921.430.35–23.0$0.3527
overallKimi K394/96 · 98%$0.006341.110.69–2.14$0.5526
overallGemini 3.6 Flash89/96 · 93%$0.006541.080.67–1.80$0.5366
overallQwen3.8 2.4T92/96 · 96%$0.007930.890.24–3.63$0.5943
overallGLM-579/96 · 82%$0.009540.740.18–4.40$0.4610
overallQwen3.8 Max90/96 · 94%$0.009590.730.23–3.38$0.6667
overallSeed 2.1 Turbo85/96 · 89%$0.01190.590.25–2.03$0.7170
overallSeed 2.0 Code85/96 · 89%$0.01460.480.24–1.35$0.9409
overallClaude Opus 5 Fast77/96 · 80%$0.01490.470.20–0.86$1.04
overallGPT-5.6 Sol Pro93/96 · 97%$0.01920.370.21–0.64$1.71
overallMuse Spark 1.20/96 · 0%unavailableunavailable$0.00e+00
verticalbankingClaude Haiku 4.518/18 · 100%$0.0002019.710.6–36.7$0.00364
verticalbankingClaude Opus 4.615/18 · 83%$0.003061.260.48–2.06$0.0530
verticalbankingClaude Sonnet 4.618/18 · 100%$0.000864.483.52–5.80$0.0158
verticalbankingDeepSeek V3.215/18 · 83%$0.000039597.440.7–234$0.00056
verticalbankingDeepSeek V4 Flash18/18 · 100%$0.000043688.247.0–159$0.00079
verticalbankingDeepSeek V4 Pro17/18 · 94%$0.0002217.410.7–28.3$0.00366
verticalbankingGemini 2.5 Flash18/18 · 100%$0.000082746.522.0–96.8$0.00160
verticalbankingGemini 3.6 Flash18/18 · 100%$0.004300.890.49–1.61$0.0786
verticalbankingGemini 3.7 Flash18/18 · 100%$0.000675.783.50–9.93$0.0124
verticalbankingGLM-4.7 Flash18/18 · 100%$0.0003012.96.98–21.7$0.00585
verticalbankingGLM-518/18 · 100%$0.001312.951.51–5.65$0.0235
verticalbankingGLM-5.218/18 · 100%$0.000954.072.02–6.82$0.0193
verticalbankingGLM-5.318/18 · 100%$0.000864.451.96–8.97$0.0175
verticalbankingGPT-5.4 mini18/18 · 100%$0.0001328.916.0–55.4$0.00246
verticalbankingGPT-5.6 Luna18/18 · 100%$0.000073452.433.3–81.4$0.00134
verticalbankingGPT-5.6 Luna Pro18/18 · 100%$0.000616.314.06–10.1$0.0110
verticalbankingGPT-5.6 Sol18/18 · 100%$0.0003311.66.36–22.8$0.00610
verticalbankingGPT-5.6 Sol Pro18/18 · 100%$0.007460.520.29–0.94$0.1358
verticalbankingGPT-5.6 Terra18/18 · 100%$0.000517.565.38–10.9$0.00905
verticalbankingGPT-5.6 Terra Pro18/18 · 100%$0.002281.681.13–2.57$0.0403
verticalbankinggpt-oss 120B18/18 · 100%$0.0000385$0.0000233–$0.0000641100$0.00069
verticalbankingGrok 4.518/18 · 100%$0.001492.591.80–3.71$0.0265
verticalbankingGrok 4.618/18 · 100%$0.002221.731.06–2.69$0.0396
verticalbankingKimi K2.518/18 · 100%$0.001213.191.83–5.56$0.0214
verticalbankingKimi K318/18 · 100%$0.004990.770.55–1.14$0.0944
verticalbankingLlama 4 Maverick9/18 · 50%$0.000448.811.96–23.2$0.00276
verticalbankingMuse Spark 1.20/18 · 0%unavailableunavailable$0.00e+00
verticalbankingClaude Opus 5 Fast16/18 · 89%$0.006390.600.31–1.02$0.1079
verticalbankingQwen3.5 Flash12/18 · 67%$0.000419.313.64–16.1$0.00530
verticalbankingQwen3.8 2.4T18/18 · 100%$0.001502.561.96–3.31$0.0265
verticalbankingQwen3.8 27B18/18 · 100%$0.000695.544.20–7.35$0.0120
verticalbankingQwen3.8 Max18/18 · 100%$0.001542.512.08–3.05$0.0270
verticalbankingSeed 2.0 Code18/18 · 100%$0.005200.740.28–2.85$0.0791
verticalbankingSeed 2.1 Turbo18/18 · 100%$0.001971.961.21–3.62$0.0324
verticalcodingClaude Haiku 4.521/21 · 100%$0.001258.016.07–11.4$0.0240
verticalcodingClaude Opus 4.621/21 · 100%$0.006011.661.36–2.37$0.1146
verticalcodingClaude Sonnet 4.621/21 · 100%$0.003352.992.32–5.01$0.0640
verticalcodingDeepSeek V3.221/21 · 100%$0.0001010068.5–173$0.00198
verticalcodingDeepSeek V4 Flash16/21 · 76%$0.004002.500.60–27.2$0.0544
verticalcodingDeepSeek V4 Pro15/21 · 71%$0.02190.460.13–5.13$0.3178
verticalcodingGemini 2.5 Flash20/21 · 95%$0.001059.576.21–17.2$0.0189
verticalcodingGemini 3.6 Flash20/21 · 95%$0.01410.710.53–1.01$0.2491
verticalcodingGemini 3.7 Flash21/21 · 100%$0.001985.044.24–6.74$0.0378
verticalcodingGLM-4.7 Flash12/21 · 57%$0.005751.740.60–5.23$0.0586
verticalcodingGLM-512/21 · 57%$0.03030.330.10–1.66$0.2887
verticalcodingGLM-5.219/21 · 90%$0.01150.870.42–3.18$0.1818
verticalcodingGLM-5.321/21 · 100%$0.005091.971.34–3.50$0.0988
verticalcodingGPT-5.4 mini21/21 · 100%$0.0007114.011.2–20.2$0.0138
verticalcodingGPT-5.6 Luna21/21 · 100%$0.0004522.113.4–48.4$0.00888
verticalcodingGPT-5.6 Luna Pro21/21 · 100%$0.002084.812.95–8.91$0.0414
verticalcodingGPT-5.6 Sol21/21 · 100%$0.003053.282.05–7.76$0.0591
verticalcodingGPT-5.6 Sol Pro21/21 · 100%$0.04800.210.13–0.51$0.9273
verticalcodingGPT-5.6 Terra21/21 · 100%$0.003892.571.56–6.06$0.0752
verticalcodingGPT-5.6 Terra Pro21/21 · 100%$0.009571.050.68–2.08$0.1874
verticalcodinggpt-oss 120B20/21 · 95%$0.00012$0.0000649–$0.0001786.9$0.00208
verticalcodingGrok 4.521/21 · 100%$0.005781.731.34–2.46$0.1104
verticalcodingGrok 4.621/21 · 100%$0.008271.210.89–1.80$0.1560
verticalcodingKimi K2.521/21 · 100%$0.008991.110.76–1.94$0.1669
verticalcodingKimi K321/21 · 100%$0.01350.740.48–1.56$0.2627
verticalcodingLlama 4 Maverick16/21 · 76%$0.0001661.833.3–101$0.00234
verticalcodingMuse Spark 1.20/21 · 0%unavailableunavailable$0.00e+00
verticalcodingClaude Opus 5 Fast20/21 · 95%$0.02400.420.29–0.63$0.4387
verticalcodingQwen3.5 Flash3/21 · 14%$0.01210.830.54–4.17$0.0210
verticalcodingQwen3.8 2.4T18/21 · 86%$0.02870.350.11–1.89$0.4290
verticalcodingQwen3.8 27B18/21 · 86%$0.01900.530.21–3.13$0.3142
verticalcodingQwen3.8 Max18/21 · 86%$0.02850.350.11–2.15$0.4268
verticalcodingSeed 2.0 Code16/21 · 76%$0.03980.250.13–0.44$0.5430
verticalcodingSeed 2.1 Turbo17/21 · 81%$0.02750.360.20–0.64$0.3939
verticalcustomer_serviceClaude Haiku 4.56/15 · 40%$0.001482.070.38–9.99$0.00609
verticalcustomer_serviceClaude Opus 4.615/15 · 100%$0.002961.040.45–1.74$0.0420
verticalcustomer_serviceClaude Sonnet 4.615/15 · 100%$0.001502.041.12–2.96$0.0214
verticalcustomer_serviceDeepSeek V3.215/15 · 100%$0.000035486.633.1–226$0.00051
verticalcustomer_serviceDeepSeek V4 Flash15/15 · 100%$0.000030710046.8–173$0.00044
verticalcustomer_serviceDeepSeek V4 Pro15/15 · 100%$0.0001916.58.30–26.7$0.00262
verticalcustomer_serviceGemini 2.5 Flash15/15 · 100%$0.000060450.820.1–94.0$0.00090
verticalcustomer_serviceGemini 3.6 Flash15/15 · 100%$0.003200.960.39–2.00$0.0462
verticalcustomer_serviceGemini 3.7 Flash15/15 · 100%$0.000515.972.35–11.2$0.00755
verticalcustomer_serviceGLM-4.7 Flash15/15 · 100%$0.0002015.06.46–27.7$0.00298
verticalcustomer_serviceGLM-514/15 · 93%$0.000943.281.04–7.91$0.0132
verticalcustomer_serviceGLM-5.215/15 · 100%$0.000634.841.67–9.46$0.00963
verticalcustomer_serviceGLM-5.313/15 · 87%$0.001182.610.60–7.51$0.0142
verticalcustomer_serviceGPT-5.4 mini12/15 · 80%$0.0001817.48.79–23.1$0.00203
verticalcustomer_serviceGPT-5.6 Luna15/15 · 100%$0.000063748.226.7–66.8$0.00094
verticalcustomer_serviceGPT-5.6 Luna Pro15/15 · 100%$0.000545.712.60–9.09$0.00793
verticalcustomer_serviceGPT-5.6 Sol15/15 · 100%$0.000496.213.55–8.94$0.00713
verticalcustomer_serviceGPT-5.6 Sol Pro15/15 · 100%$0.009860.310.14–0.47$0.1461
verticalcustomer_serviceGPT-5.6 Terra15/15 · 100%$0.000565.493.33–7.56$0.00808
verticalcustomer_serviceGPT-5.6 Terra Pro15/15 · 100%$0.002101.460.70–2.19$0.0308
verticalcustomer_servicegpt-oss 120B15/15 · 100%$0.0000511$0.0000189–$0.0001160.1$0.00072
verticalcustomer_serviceGrok 4.515/15 · 100%$0.001671.840.89–3.07$0.0243
verticalcustomer_serviceGrok 4.615/15 · 100%$0.002561.200.62–1.99$0.0369
verticalcustomer_serviceKimi K2.515/15 · 100%$0.001162.651.51–3.91$0.0166
verticalcustomer_serviceKimi K315/15 · 100%$0.002341.310.51–2.60$0.0347
verticalcustomer_serviceLlama 4 Maverick11/15 · 73%$0.0002512.52.58–32.5$0.00245
verticalcustomer_serviceMuse Spark 1.20/15 · 0%unavailableunavailable$0.00e+00
verticalcustomer_serviceClaude Opus 5 Fast12/15 · 80%$0.01110.280.04–1.19$0.1139
verticalcustomer_serviceQwen3.5 Flash9/15 · 60%$0.000644.821.41–8.74$0.00582
verticalcustomer_serviceQwen3.8 2.4T15/15 · 100%$0.001342.291.21–3.63$0.0190
verticalcustomer_serviceQwen3.8 27B15/15 · 100%$0.000585.312.66–8.66$0.00803
verticalcustomer_serviceQwen3.8 Max15/15 · 100%$0.001392.211.34–3.20$0.0197
verticalcustomer_serviceSeed 2.0 Code15/15 · 100%$0.003510.870.57–2.22$0.0490
verticalcustomer_serviceSeed 2.1 Turbo15/15 · 100%$0.002581.190.67–3.64$0.0363
verticallegalClaude Haiku 4.56/15 · 40%$0.0007715.25.01–28.9$0.00448
verticallegalClaude Opus 4.614/15 · 93%$0.002794.212.24–7.70$0.0420
verticallegalClaude Sonnet 4.610/15 · 67%$0.002574.581.48–8.92$0.0284
verticallegalDeepSeek V3.26/15 · 40%$0.0001577.926.0–153$0.00094
verticallegalDeepSeek V4 Flash9/15 · 60%$0.0001576.134.6–137$0.00147
verticallegalDeepSeek V4 Pro12/15 · 80%$0.002205.362.17–14.2$0.0248
verticallegalGemini 2.5 Flash6/15 · 40%$0.0002941.113.6–75.5$0.00169
verticallegalGemini 3.6 Flash9/15 · 60%$0.008001.470.42–2.87$0.0815
verticallegalGemini 3.7 Flash11/15 · 73%$0.001279.283.95–19.1$0.0142
verticallegalGLM-4.7 Flash4/15 · 27%$0.001468.062.28–19.8$0.00618
verticallegalGLM-58/15 · 53%$0.01410.840.14–5.91$0.1053
verticallegalGLM-5.212/15 · 80%$0.001906.194.28–9.24$0.0226
verticallegalGLM-5.39/15 · 60%$0.003523.340.81–7.79$0.0282
verticallegalGPT-5.4 mini4/15 · 27%$0.0009312.63.24–28.9$0.00325
verticallegalGPT-5.6 Luna13/15 · 87%$0.0001577.941.5–146$0.00189
verticallegalGPT-5.6 Luna Pro15/15 · 100%$0.0009112.96.41–24.2$0.0133
verticallegalGPT-5.6 Sol12/15 · 80%$0.0011210.56.60–17.5$0.0133
verticallegalGPT-5.6 Sol Pro12/15 · 80%$0.02080.570.30–0.98$0.2430
verticallegalGPT-5.6 Terra12/15 · 80%$0.001229.615.12–15.8$0.0144
verticallegalGPT-5.6 Terra Pro12/15 · 80%$0.004172.821.35–4.94$0.0488
verticallegalgpt-oss 120B12/15 · 80%$0.00012$0.0000476–$0.00024100$0.00139
verticallegalGrok 4.514/15 · 93%$0.003473.392.07–5.61$0.0469
verticallegalGrok 4.613/15 · 87%$0.006891.711.03–3.02$0.0866
verticallegalKimi K2.512/15 · 80%$0.004942.381.36–4.34$0.0562
verticallegalKimi K313/15 · 87%$0.008841.330.52–3.67$0.1051
verticallegalLlama 4 Maverick7/15 · 47%$0.0004924.05.54–72.3$0.00277
verticallegalMuse Spark 1.20/15 · 0%unavailableunavailable$0.00e+00
verticallegalClaude Opus 5 Fast11/15 · 73%$0.02310.510.14–1.59$0.2177
verticallegalQwen3.5 Flash8/15 · 53%$0.001557.593.28–14.8$0.0127
verticallegalQwen3.8 2.4T14/15 · 93%$0.006921.700.68–5.22$0.0901
verticallegalQwen3.8 27B12/15 · 80%$0.002654.442.49–10.8$0.0297
verticallegalQwen3.8 Max12/15 · 80%$0.01540.760.32–5.28$0.1647
verticallegalSeed 2.0 Code10/15 · 67%$0.02140.550.20–1.57$0.2039
verticallegalSeed 2.1 Turbo8/15 · 53%$0.02540.460.13–1.68$0.2045
verticalmedicalClaude Haiku 4.524/27 · 89%$0.000436.463.60–14.8$0.00925
verticalmedicalClaude Opus 4.627/27 · 100%$0.001951.410.78–3.43$0.0505
verticalmedicalClaude Sonnet 4.627/27 · 100%$0.001122.461.66–4.02$0.0280
verticalmedicalDeepSeek V3.226/27 · 96%$0.000040168.548.3–104$0.00098
verticalmedicalDeepSeek V4 Flash27/27 · 100%$0.000027510079.3–126$0.00072
verticalmedicalDeepSeek V4 Pro26/27 · 96%$0.0001517.912.7–24.8$0.00379
verticalmedicalGemini 2.5 Flash25/27 · 93%$0.0001320.610.2–57.2$0.00309
verticalmedicalGemini 3.6 Flash27/27 · 100%$0.003130.880.66–1.22$0.0813
verticalmedicalGemini 3.7 Flash27/27 · 100%$0.000515.434.14–6.88$0.0134
verticalmedicalGLM-4.7 Flash25/27 · 93%$0.000397.032.46–14.1$0.0124
verticalmedicalGLM-527/27 · 100%$0.001142.421.79–3.64$0.0303
verticalmedicalGLM-5.227/27 · 100%$0.000684.052.81–6.24$0.0186
verticalmedicalGLM-5.324/27 · 89%$0.000693.992.21–6.18$0.0160
verticalmedicalGPT-5.4 mini24/27 · 89%$0.0001518.411.9–24.8$0.00350
verticalmedicalGPT-5.6 Luna27/27 · 100%$0.000075736.328.4–46.3$0.00195
verticalmedicalGPT-5.6 Luna Pro27/27 · 100%$0.000584.713.47–6.26$0.0151
verticalmedicalGPT-5.6 Sol27/27 · 100%$0.000485.694.65–6.99$0.0123
verticalmedicalGPT-5.6 Sol Pro27/27 · 100%$0.009890.280.22–0.34$0.2547
verticalmedicalGPT-5.6 Terra27/27 · 100%$0.000634.383.36–5.65$0.0160
verticalmedicalGPT-5.6 Terra Pro27/27 · 100%$0.002641.040.77–1.40$0.0675
verticalmedicalgpt-oss 120B27/27 · 100%$0.0000297$0.0000202–$0.000042392.6$0.00076
verticalmedicalGrok 4.527/27 · 100%$0.001721.601.29–1.97$0.0441
verticalmedicalGrok 4.627/27 · 100%$0.002011.371.13–1.63$0.0521
verticalmedicalKimi K2.527/27 · 100%$0.001322.091.17–3.88$0.0323
verticalmedicalKimi K327/27 · 100%$0.002031.351.02–1.65$0.0558
verticalmedicalLlama 4 Maverick13/27 · 48%$0.000348.002.00–24.7$0.00336
verticalmedicalMuse Spark 1.20/27 · 0%unavailableunavailable$0.00e+00
verticalmedicalClaude Opus 5 Fast18/27 · 67%$0.01010.270.08–0.63$0.1625
verticalmedicalQwen3.5 Flash21/27 · 78%$0.000446.282.73–15.7$0.00891
verticalmedicalQwen3.8 2.4T27/27 · 100%$0.001162.381.87–2.95$0.0298
verticalmedicalQwen3.8 27B27/27 · 100%$0.000515.374.42–6.35$0.0135
verticalmedicalQwen3.8 Max27/27 · 100%$0.001092.531.94–3.18$0.0285
verticalmedicalSeed 2.0 Code26/27 · 96%$0.002850.960.47–2.57$0.0658
verticalmedicalSeed 2.1 Turbo27/27 · 100%$0.002091.320.71–3.19$0.0500