BEP Research Research preview

20 models, priced per finished task.

The current board — every model run on the same 32 tasks, graded the same way, and costed on what it actually spent to get work accepted. The 6 models at the top are doing work we cannot tell apart, at a 73× spread in price.

Price is what you pay for tokens. Cost is what you pay for mistakes.

benchmark 0.3.0 · pricing 2026-07-09 · billed cost · baseline gpt-oss 120B · 32 tasks · computed 2026-08-04 17:11
20 AI models tested1,920 graded testsAugust 2026 last updated
What are you trying to automate?
How much of it, and what does a mistake cost?
Tasks per month
A wrong answer that reaches a customer costs us
Most teams underestimate this. $200 is a reasonable starting point for a customer-facing mistake that needs finding, apologising for and redoing.
Accuracy we require

Workload calculator

The benchmark measures what an attempt costs and how often it is accepted. Everything else that decides your bill is yours: volume, review time, what a wrong answer costs when it reaches a customer. Supply those and the ranking can change — which is the point.

Retries assume a failed task fails again, because that is what we measured — asking a second time usually returns the same wrong answer at twice the price. Why that is, and when it is not true →

Your workload assumption
Human review assumption
Risk and retries assumption

Success rates and cost per attempt are measured on the benchmark tasks shown below. Volume, review time, error cost, detection rate and retry behaviour are assumptions you supplied. Totals combine both and are estimates, not quotes.

The evidence

Everything below is how that recommendation was reached: what each model cost, how often a grader accepted its output, and where the numbers are too thin to act on. The Octane score explains the ranking; it is not the product.

$0.68 per 10,000 completed tasks · gpt-oss 120B · 155.3× cheaper than GLM-5
baseline
observed success of that option
95%
521 of 10,000 still need redoing
weakest vertical
80%
Legal — weakest measured vertical
widest gap in one vertical
417×
how much the choice is worth
kinds of work tested
5
1,920 graded tests

Which model won each kind of work

These cards reflect performance on the specific benchmark tasks listed under “What we test” — they are starting points for a model evaluation, not blanket recommendations for an entire industry. The badge is the observed success rate on those tasks, with the run counts beneath it; not a claim about production reliability. Each card leads with the highest-scoring model above the 80% threshold, then names a lower-cost option that also cleared it. Where no model reached the threshold, the card says so instead of recommending one.

Banking & Finance
100%observed success
18 / 18 benchmark runs
DeepSeek V3.2
$0.32 per 10,000 successful tasks
Lowest measured cost above the threshold
Customer Service
100%observed success
15 / 15 benchmark runs
gpt-oss 120B
$0.30 per 10,000 successful tasks
Lowest measured cost above the threshold
Legal
100%observed success
15 / 15 benchmark runs
GPT-5.6 Luna Pro
$4.36 per 10,000 successful tasks
Lower-cost option: GPT-5.6 Luna — 14/15 (93%), $0.76, 5.7× cheaper
Medical
100%observed success
27 / 27 benchmark runs
DeepSeek V4 Flash
$0.27 per 10,000 successful tasks
Lowest measured cost above the threshold
Software Engineering
100%observed success
21 / 21 benchmark runs
DeepSeek V3.2
$0.98 per 10,000 successful tasks
Lowest measured cost above the threshold

Sticker price is not the price

A pricing page sells you a rate per million tokens. That rate is one of three terms, and it is the only one published. A model can advertise the lowest rate, write two or three times as many tokens to answer the same question, fail a share of the time — and finish as the most expensive option here.

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

what the price page implies →$ per 10,000 completed tasksgpt-oss 120B$0.12$0.56$0.59$0.68×4.9 tokens · ×1.1 failures · ×1.1 work mix → ×5.8 the stickerDeepSeek V3.2$0.56$0.48$0.57$0.72×0.9 tokens · ×1.2 failures · ×1.3 work mix → ×1.3 the stickerGPT-5.6 Luna$0.36$0.75$0.76$0.77×2.1 tokens · ×1.0 failures · ×1.0 work mix → ×2.1 the stickerGemini 2.5 Flash$1.32$2.56$2.82$2.89×1.9 tokens · ×1.1 failures · ×1.0 work mix → ×2.2 the stickerLlama 4 Maverick$0.59$1.43$2.75$4.07×2.4 tokens · ×1.9 failures · ×1.5 work mix → ×7.0 the stickerDeepSeek V4 Flash$0.20$2.89$3.30$4.34×14.2 tokens · ×1.1 failures · ×1.3 work mix → ×21.4 the stickerGPT-5.4 mini$2.70$2.60$3.11$4.34×1.0 tokens · ×1.2 failures · ×1.4 work mix → ×1.6 the stickerGPT-5.6 Luna Pro$0.36$4.58$4.58$4.67×12.7 tokens · ×1.0 failures · ×1.0 work mix → ×13.0 the stickerGPT-5.6 Terra$3.61$6.44$6.65$6.84×1.8 tokens · ×1.0 failures · ×1.0 work mix → ×1.9 the stickerClaude Haiku 4.5$3.27$4.85$6.21$8.13×1.5 tokens · ×1.3 failures · ×1.3 work mix → ×2.5 the stickerGLM-4.7 Flash$0.23$9.21$11.48$14.54×40.1 tokens · ×1.2 failures · ×1.3 work mix → ×63.3 the stickerClaude Sonnet 4.6$9.80$16.77$17.49$19.00×1.7 tokens · ×1.0 failures · ×1.1 work mix → ×1.9 the stickerGLM-5.2$3.14$21.64$22.34$23.31×6.9 tokens · ×1.0 failures · ×1.0 work mix → ×7.4 the stickerQwen3.5 Flash$0.19$5.35$9.70$29.05×28.1 tokens · ×1.8 failures · ×3.0 work mix → ×152.5 the stickerClaude Opus 4.6$16.34$31.30$32.66$33.54×1.9 tokens · ×1.0 failures · ×1.0 work mix → ×2.1 the stickerGPT-5.6 Sol$18.03$32.08$33.12$34.20×1.8 tokens · ×1.0 failures · ×1.0 work mix → ×1.9 the stickerKimi K2.5$1.86$35.02$37.35$43.80×18.8 tokens · ×1.1 failures · ×1.2 work mix → ×23.5 the stickerKimi K3$9.80$53.60$54.16$56.46×5.5 tokens · ×1.0 failures · ×1.0 work mix → ×5.8 the stickerQwen3.8 Max$5.19$67.35$71.05$89.88×13.0 tokens · ×1.1 failures · ×1.3 work mix → ×17.3 the stickerGLM-5$2.36$55.08$66.09$105.13×23.3 tokens · ×1.2 failures · ×1.6 work mix → ×44.5 the sticker$0$28$55$83$110

Each bar adds one thing the price page does not tell you. The faintest is the advertised rate on this exact workload, priced as if the model were as concise as the leanest one measured. Then the tokens it really wrote. Then the attempts you paid for and could not use. Then your actual work mix, since the index weights every vertical equally instead of letting a model's weak vertical be diluted by its strong one. Only the faintest bar is visible when you pick a model.

Overall

The leading model in each panel sets 100; every other grade is its share of that leader's successful work per dollar. Log axis — a real panel spans orders of magnitude. Whiskers are 95% clustered bootstrap intervals, resampled by vertical, then task, then repetition. Raw baseline-relative scores are in the CSV and JSON exports.

gpt-oss 120BClaude Haiku 4.5Claude Sonnet 4.6Claude Opus 4.6Qwen3.5 FlashLlama 4 MaverickDeepSeek V3.2Kimi K2.5GLM-5GLM-5.2GLM-4.7 FlashGPT-5.4 miniGPT-5.6 LunaGPT-5.6 Luna ProGPT-5.6 TerraGPT-5.6 SolGemini 2.5 FlashDeepSeek V4 FlashKimi K3Qwen3.8 Max
0.010.101.0010.01001,000100 = scope leadergpt-oss 120B100DeepSeek V3.294.6GPT-5.6 Luna87.4Gemini 2.5 Flash23.4Llama 4 Maverick16.650/96 — below thresholdDeepSeek V4 Flash15.6GPT-5.4 mini15.6GPT-5.6 Luna Pro14.5GPT-5.6 Terra9.89Claude Haiku 4.58.3275/96 — below thresholdGLM-4.7 Flash4.66Claude Sonnet 4.63.56GLM-5.22.90Qwen3.5 Flash2.3353/96 — below thresholdClaude Opus 4.62.02GPT-5.6 Sol1.98Kimi K2.51.55Kimi K31.20Qwen3.8 Max0.75GLM-50.64
Why the overall leader differs from the vertical leaders

gpt-oss 120B ranks first overall because its lead in Customer Service outweighs DeepSeek V3.2's lead in Banking & Finance, Software Engineering under the benchmark's equal-weight methodology — each vertical counts the same regardless of how many tasks or attempts it holds. Weight the verticals to your own workload in the calculator above and the ranking can change.

Observations
  • The choice of model matters most in Software Engineering, where measured cost per successful task spans 417x.
  • GPT-5.6 Luna Pro completed every benchmark run (96/96).

By vertical

Every panel shares one axis, so they are directly comparable. This is the cut that matters — cost-effectiveness is a property of a model on a kind of work, not of a model alone.

A hollow, struck-through dot marks a model that recorded below 80% observed success on that vertical. It is excluded from the recommendation, not judged unusable — cost per successful task already prices retries, but it assumes a failure is free to detect and that retrying eventually works. Neither held in these measurements: most model-task pairs returned byte-identical text on every repetition, so a failed attempt repeated rather than varied.

Banking & Finance
0.010.101.0010.01001,000DeepSeek V3.2100GPT-5.6 Luna91.5DeepSeek V4 Flash90.1gpt-oss 120B76.6Gemini 2.5 Flash39.0GPT-5.4 mini24.6Claude Haiku 4.516.5GPT-5.6 Terra12.6GPT-5.6 Luna Pro10.8Llama 4 Maverick7.34Qwen3.5 Flash6.76GLM-4.7 Flash6.28Claude Sonnet 4.63.75GPT-5.6 Sol3.58GLM-5.23.46Kimi K2.52.20Qwen3.8 Max2.14GLM-51.57Claude Opus 4.61.12Kimi K30.89
  • DeepSeek V3.2 recorded the lowest measured cost among models above the threshold, at 18/18 successful · 100% observed.
  • Measured cost per successful task spans 113x across the field, from DeepSeek V3.2 to Kimi K3.
Customer Service
0.010.101.0010.01001,000gpt-oss 120B100GPT-5.6 Luna93.8DeepSeek V3.280.6DeepSeek V4 Flash78.2Gemini 2.5 Flash49.7GPT-5.4 mini18.3GLM-4.7 Flash14.8GPT-5.6 Luna Pro12.0GPT-5.6 Terra9.94Llama 4 Maverick7.23Qwen3.5 Flash6.24GLM-5.24.19Qwen3.8 Max2.34GPT-5.6 Sol2.16Claude Haiku 4.52.05Claude Sonnet 4.61.95Kimi K2.51.77GLM-51.68Kimi K31.30Claude Opus 4.61.03
  • gpt-oss 120B recorded the lowest measured cost among models above the threshold, at 15/15 successful · 100% observed.
  • Measured cost per successful task spans 97x across the field, from gpt-oss 120B to Claude Opus 4.6.
Legal
0.010.101.0010.01001,000GPT-5.6 Luna100gpt-oss 120B74.2DeepSeek V3.250.6DeepSeek V4 Flash26.8Gemini 2.5 Flash26.6GPT-5.6 Luna Pro17.4GPT-5.6 Terra12.7Llama 4 Maverick12.1Claude Haiku 4.59.84GPT-5.4 mini7.49Qwen3.5 Flash4.67GLM-4.7 Flash4.42GLM-5.24.07Claude Sonnet 4.62.96Claude Opus 4.62.39GPT-5.6 Sol2.33GLM-51.24Kimi K2.51.12Kimi K30.92Qwen3.8 Max0.64
  • GPT-5.6 Luna recorded the lowest measured cost among models above the threshold, at 14/15 successful · 93% observed.
  • Measured cost per successful task spans 156x across the field, from GPT-5.6 Luna to Qwen3.8 Max.
Medical
0.010.101.0010.01001,000DeepSeek V4 Flash100gpt-oss 120B92.8GPT-5.6 Luna76.1DeepSeek V3.267.3Gemini 2.5 Flash22.4GPT-5.4 mini18.2GLM-4.7 Flash10.9GPT-5.6 Luna Pro9.18GPT-5.6 Terra9.09Qwen3.5 Flash8.64Llama 4 Maverick6.86Claude Haiku 4.56.44GLM-5.23.99Claude Sonnet 4.62.37Qwen3.8 Max2.37GPT-5.6 Sol2.07Kimi K2.51.60GLM-51.53Claude Opus 4.61.42Kimi K31.39
  • DeepSeek V4 Flash recorded the lowest measured cost among models above the threshold, at 27/27 successful · 100% observed.
  • Measured cost per successful task spans 72x across the field, from DeepSeek V4 Flash to Kimi K3.
Software Engineering
0.010.101.0010.01001,000DeepSeek V3.2100gpt-oss 120B72.6Llama 4 Maverick62.1GPT-5.6 Luna47.0GPT-5.4 mini13.8Gemini 2.5 Flash10.9GPT-5.6 Luna Pro9.26Claude Haiku 4.58.08DeepSeek V4 Flash5.48GPT-5.6 Terra4.97Claude Sonnet 4.62.88GLM-4.7 Flash2.13Claude Opus 4.61.66GLM-5.21.31Kimi K2.50.96GPT-5.6 Sol0.96Qwen3.5 Flash0.84Kimi K30.81Qwen3.8 Max0.34GLM-50.24
  • DeepSeek V3.2 recorded the lowest measured cost among models above the threshold, at 21/21 successful · 100% observed.
  • Measured cost per successful task spans 417x across the field, from DeepSeek V3.2 to GLM-5.

The frontier

Observed success against cost per successful task. Bottom-right is the good corner: high observed success at low cost. Position carries the reading, not colour — every point is directly labelled.

$0.00001$0.00010$0.00100$0.0100$0.10000%25%50%75%100%success rate → bettercost per success ↓ bettergpt-oss 120BDeepSeek V3.2GPT-5.6 LunaGemini 2.5 FlashLlama 4 MaverickDeepSeek V4 FlashGPT-5.4 miniGPT-5.6 Luna ProGPT-5.6 TerraClaude Haiku 4.5GLM-4.7 FlashClaude Sonnet 4.6GLM-5.2Qwen3.5 FlashClaude Opus 4.6GPT-5.6 SolKimi K2.5Kimi K3Qwen3.8 MaxGLM-5

Methodology

Cost per Successful Task = Total Billed API Cost Across All Attempts ÷ Successful Outputs

Total billed cost includes input tokens, output tokens, cached tokens where applicable, reasoning tokens where separately billed, failed attempts, and retries where benchmarked. Work you paid for and could not use stays in the numerator.

Full methodology, grader definitions, confidence rules and limitations →

What we test

Every task behind these numbers, with how it is graded and what it weighs. Full prompts are deliberately unpublished — printed prompts end up in training data and the benchmark decays — but per-task results are in the CSV and JSON exports, so any figure on this page traces to graded attempts.

Banking & Finance 6 tasks

TaskGraded byWeightVersion
Accrue interest under a stated day-count conventionexact answer2v1
Convert a nominal rate to an effective annual ratestructured output, field-checked1.5v1
Refuse to settle an FX conversion at the wrong date's rateexact answer2v1
Assign a KYC risk tierexact answer1v1
Allocate a partial payment down a contractual waterfallstructured output, field-checked1.5v1
Extract structured fields from a bank statement linestructured output, field-checked1v1

Customer Service 5 tasks

TaskGraded byWeightVersion
Apply a return policy where the exception overrides the deadlineexact answer2v1
Apply a multi-condition escalation ruleexact answer1.5v1
Compute a prorated refund on cancellationstructured output, field-checked1.5v1
Decide refund eligibility against a policystructured output, field-checked1v1
Classify a support ticket into a routing intentexact answer1v1

Legal 5 tasks

TaskGraded byWeightVersion
Resolve conflicting governing-law clauses via a precedence rulestructured output, field-checked1.5v1
Extract governing law and venue from a contract clausestructured output, field-checked1v1
Apply a liability cap with a carve-out and an absolute exclusionexact answer2v1
Compute a notice expiry with exclusion and weekend roll-forwardstructured output, field-checked1.5v1
Earliest termination date under an initial-term barexact answer2v1

Medical 9 tasks

TaskGraded byWeightVersion
Apply a drug interaction only when its stated condition holdsexact answer2v1
Identify the ICD-10 code for a described conditionexact answer1v2
Convert a weight-based infusion order into a pump rateexact answer2v1
Extract a drug interaction into structured formstructured output, field-checked1v1
Apply a maximum-dose ceiling before converting to volumeexact answer2v1
Refuse to fill a weight-based order when the weight is missingexact answer2v1
Convert a weight-based dose into a dispensing volumestructured output, field-checked1.5v1
Interpret a lab value against a sex-specific reference rangeexact answer1.5v1
Adjust a dose for renal function using Cockcroft-Gaultexact answer2v1

Software Engineering 7 tasks

TaskGraded byWeightVersion
Add months to a date, clamping to the end of the monthcode executed against tests2v1
Implement order-preserving deduplicationcode executed against tests1.5v1
Predict the output of an order-preserving deduplicationexact answer1v1
Merge overlapping intervals, including touching endpointscode executed against tests1.5v1
Count divisible pairs at a scale that punishes brute forcecode executed against tests2v1
Parse durations, honouring a buried error-handling contractcode executed against tests1.5v1
Subtract one set of half-open intervals from anothercode executed against tests2v1

Benchmark your own workload

This public basket is 23 short, single-shot tasks. Your work is not these tasks. The same machinery runs against representative samples of your own workload — extraction from your documents, your support tickets, your code — and returns model comparison, total-cost analysis under your review and error costs, observed success by task type, routing recommendations, and the caveats that apply.

Useful for model selection, vendor evaluation, cost optimisation, routing design and procurement support. Get in touch to benchmark your workload →

All figures

Task-level rows are omitted here for length. The CSV and JSON exports carry them in full.

ScopeVerticalModelObserved successCost per successful taskCost 95% CIOctaneOctane 95% CITotal billed cost
overallgpt-oss 120B91/96 · 95%$0.0000677$0.0000299–$0.00013100$0.00540
overallDeepSeek V3.281/96 · 84%$0.000071694.629.5–178$0.00460
overallGPT-5.6 Luna95/96 · 99%$0.000077587.454.5–147$0.00718
overallGemini 2.5 Flash87/96 · 91%$0.0002923.414.0–56.0$0.0246
overallLlama 4 Maverick50/96 · 52%$0.0004116.64.27–41.6$0.0137
overallDeepSeek V4 Flash84/96 · 88%$0.0004315.64.91–108$0.0277
overallGPT-5.4 mini80/96 · 83%$0.0004315.67.07–34.9$0.0249
overallGPT-5.6 Luna Pro96/96 · 100%$0.0004714.59.71–24.0$0.0439
overallGPT-5.6 Terra93/96 · 97%$0.000689.895.78–17.6$0.0618
overallClaude Haiku 4.575/96 · 78%$0.000818.322.60–16.6$0.0466
overallGLM-4.7 Flash77/96 · 80%$0.001454.661.63–15.9$0.0884
overallClaude Sonnet 4.692/96 · 96%$0.001903.561.97–5.74$0.1609
overallGLM-5.293/96 · 97%$0.002332.901.51–5.70$0.2078
overallQwen3.5 Flash53/96 · 55%$0.002912.331.30–20.1$0.0514
overallClaude Opus 4.692/96 · 96%$0.003352.021.11–3.44$0.3005
overallGPT-5.6 Sol93/96 · 97%$0.003421.981.11–3.82$0.3080
overallKimi K2.590/96 · 94%$0.004381.551.02–2.72$0.3362
overallKimi K395/96 · 99%$0.005651.200.78–1.88$0.5145
overallQwen3.8 Max91/96 · 95%$0.008990.750.26–2.90$0.6465
overallGLM-580/96 · 83%$0.01050.640.20–2.39$0.5287
verticalbankingClaude Haiku 4.518/18 · 100%$0.0002016.58.94–29.1$0.00364
verticalbankingClaude Opus 4.615/18 · 83%$0.002891.120.46–1.83$0.0496
verticalbankingClaude Sonnet 4.618/18 · 100%$0.000863.752.94–4.97$0.0158
verticalbankingDeepSeek V3.218/18 · 100%$0.000032210055.6–190$0.00058
verticalbankingDeepSeek V4 Flash18/18 · 100%$0.000035890.152.6–142$0.00068
verticalbankingGemini 2.5 Flash18/18 · 100%$0.000082739.018.9–77.6$0.00160
verticalbankingGLM-4.7 Flash17/18 · 94%$0.000516.281.67–17.0$0.0111
verticalbankingGLM-518/18 · 100%$0.002051.570.93–2.64$0.0373
verticalbankingGLM-5.218/18 · 100%$0.000933.461.99–5.19$0.0182
verticalbankingGPT-5.4 mini18/18 · 100%$0.0001324.613.7–43.3$0.00244
verticalbankingGPT-5.6 Luna18/18 · 100%$0.000035291.558.0–137$0.00065
verticalbankingGPT-5.6 Luna Pro18/18 · 100%$0.0003010.86.99–16.7$0.00539
verticalbankingGPT-5.6 Sol18/18 · 100%$0.000903.581.95–6.49$0.0167
verticalbankingGPT-5.6 Terra18/18 · 100%$0.0002612.69.22–17.2$0.00454
verticalbankinggpt-oss 120B18/18 · 100%$0.0000421$0.0000262–$0.000068976.6$0.00075
verticalbankingKimi K2.518/18 · 100%$0.001472.201.11–4.00$0.0277
verticalbankingKimi K318/18 · 100%$0.003640.890.60–1.27$0.0684
verticalbankingLlama 4 Maverick9/18 · 50%$0.000447.341.58–22.8$0.00273
verticalbankingQwen3.5 Flash11/18 · 61%$0.000486.762.09–12.8$0.00586
verticalbankingQwen3.8 Max18/18 · 100%$0.001512.141.70–2.73$0.0266
verticalcodingClaude Haiku 4.521/21 · 100%$0.001218.085.39–13.7$0.0233
verticalcodingClaude Opus 4.621/21 · 100%$0.005881.661.14–2.80$0.1123
verticalcodingClaude Sonnet 4.621/21 · 100%$0.003392.881.99–5.55$0.0647
verticalcodingDeepSeek V3.221/21 · 100%$0.000097810067.8–194$0.00189
verticalcodingDeepSeek V4 Flash15/21 · 71%$0.001785.481.81–25.0$0.0231
verticalcodingGemini 2.5 Flash21/21 · 100%$0.0009010.96.80–18.8$0.0173
verticalcodingGLM-4.7 Flash14/21 · 67%$0.004582.130.67–7.90$0.0575
verticalcodingGLM-511/21 · 52%$0.04080.240.08–0.64$0.3577
verticalcodingGLM-5.221/21 · 100%$0.007461.310.73–3.45$0.1374
verticalcodingGPT-5.4 mini21/21 · 100%$0.0007113.89.18–23.3$0.0137
verticalcodingGPT-5.6 Luna21/21 · 100%$0.0002147.024.0–98.9$0.00413
verticalcodingGPT-5.6 Luna Pro21/21 · 100%$0.001069.265.39–18.3$0.0208
verticalcodingGPT-5.6 Sol21/21 · 100%$0.01020.960.53–2.74$0.1991
verticalcodingGPT-5.6 Terra21/21 · 100%$0.001974.972.83–12.6$0.0383
verticalcodinggpt-oss 120B19/21 · 90%$0.00013$0.0000673–$0.0002372.6$0.00225
verticalcodingKimi K2.521/21 · 100%$0.01020.960.52–2.20$0.1844
verticalcodingKimi K321/21 · 100%$0.01210.810.36–1.89$0.2484
verticalcodingLlama 4 Maverick16/21 · 76%$0.0001662.129.1–115$0.00236
verticalcodingQwen3.5 Flash3/21 · 14%$0.01160.840.48–4.42$0.0202
verticalcodingQwen3.8 Max18/21 · 86%$0.02910.340.10–2.22$0.4374
verticalcustomer_serviceClaude Haiku 4.56/15 · 40%$0.001472.050.65–6.90$0.00604
verticalcustomer_serviceClaude Opus 4.615/15 · 100%$0.002901.030.71–2.60$0.0412
verticalcustomer_serviceClaude Sonnet 4.615/15 · 100%$0.001541.951.39–4.13$0.0219
verticalcustomer_serviceDeepSeek V3.215/15 · 100%$0.000037280.661.5–122$0.00054
verticalcustomer_serviceDeepSeek V4 Flash15/15 · 100%$0.000038378.268.4–92.4$0.00054
verticalcustomer_serviceGemini 2.5 Flash15/15 · 100%$0.000060449.736.8–59.0$0.00090
verticalcustomer_serviceGLM-4.7 Flash15/15 · 100%$0.0002014.812.9–17.3$0.00294
verticalcustomer_serviceGLM-515/15 · 100%$0.001791.681.24–2.08$0.0259
verticalcustomer_serviceGLM-5.215/15 · 100%$0.000724.192.99–5.41$0.0111
verticalcustomer_serviceGPT-5.4 mini13/15 · 87%$0.0001618.313.1–21.4$0.00206
verticalcustomer_serviceGPT-5.6 Luna15/15 · 100%$0.00003293.866.8–143$0.00047
verticalcustomer_serviceGPT-5.6 Luna Pro15/15 · 100%$0.0002512.09.04–16.9$0.00369
verticalcustomer_serviceGPT-5.6 Sol15/15 · 100%$0.001392.161.50–3.13$0.0201
verticalcustomer_serviceGPT-5.6 Terra15/15 · 100%$0.000309.946.71–15.7$0.00435
verticalcustomer_servicegpt-oss 120B15/15 · 100%$0.00003$0.0000222–$0.0000387100$0.00043
verticalcustomer_serviceKimi K2.515/15 · 100%$0.001691.771.19–2.76$0.0243
verticalcustomer_serviceKimi K315/15 · 100%$0.002311.301.07–1.61$0.0334
verticalcustomer_serviceLlama 4 Maverick9/15 · 60%$0.000417.231.57–18.6$0.00283
verticalcustomer_serviceQwen3.5 Flash9/15 · 60%$0.000486.241.46–17.0$0.00442
verticalcustomer_serviceQwen3.8 Max15/15 · 100%$0.001282.341.83–2.95$0.0183
verticallegalClaude Haiku 4.56/15 · 40%$0.000779.843.68–15.3$0.00448
verticallegalClaude Opus 4.614/15 · 93%$0.003182.391.32–5.20$0.0480
verticallegalClaude Sonnet 4.611/15 · 73%$0.002572.960.99–6.59$0.0299
verticallegalDeepSeek V3.25/15 · 33%$0.0001550.69.29–78.8$0.00075
verticallegalDeepSeek V4 Flash9/15 · 60%$0.0002826.89.48–76.1$0.00269
verticallegalGemini 2.5 Flash6/15 · 40%$0.0002926.610.1–42.9$0.00169
verticallegalGLM-4.7 Flash6/15 · 40%$0.001724.420.80–16.6$0.0107
verticallegalGLM-510/15 · 67%$0.006121.240.56–2.52$0.0634
verticallegalGLM-5.212/15 · 80%$0.001874.072.75–5.20$0.0223
verticallegalGPT-5.4 mini4/15 · 27%$0.001027.492.64–15.2$0.00322
verticallegalGPT-5.6 Luna14/15 · 93%$0.00007610069.2–155$0.00101
verticallegalGPT-5.6 Luna Pro15/15 · 100%$0.0004417.410.5–28.6$0.00629
verticallegalGPT-5.6 Sol12/15 · 80%$0.003262.331.77–2.96$0.0385
verticallegalGPT-5.6 Terra12/15 · 80%$0.0006012.78.21–16.8$0.00704
verticallegalgpt-oss 120B12/15 · 80%$0.00010$0.0000486–$0.0001974.2$0.00121
verticallegalKimi K2.510/15 · 67%$0.006811.120.70–2.99$0.0603
verticallegalKimi K314/15 · 93%$0.008250.920.56–2.02$0.1109
verticallegalLlama 4 Maverick4/15 · 27%$0.0006312.13.03–23.7$0.00237
verticallegalQwen3.5 Flash9/15 · 60%$0.001634.672.37–8.17$0.0143
verticallegalQwen3.8 Max13/15 · 87%$0.01190.640.30–3.03$0.1341
verticalmedicalClaude Haiku 4.524/27 · 89%$0.000426.443.68–13.6$0.00919
verticalmedicalClaude Opus 4.627/27 · 100%$0.001911.420.82–3.17$0.0493
verticalmedicalClaude Sonnet 4.627/27 · 100%$0.001152.371.70–4.21$0.0286
verticalmedicalDeepSeek V3.222/27 · 81%$0.000040467.346.0–89.3$0.00084
verticalmedicalDeepSeek V4 Flash27/27 · 100%$0.000027210077.9–125$0.00072
verticalmedicalGemini 2.5 Flash27/27 · 100%$0.0001222.411.3–59.7$0.00309
verticalmedicalGLM-4.7 Flash25/27 · 93%$0.0002510.98.35–12.6$0.00625
verticalmedicalGLM-526/27 · 96%$0.001781.531.26–1.88$0.0445
verticalmedicalGLM-5.227/27 · 100%$0.000683.993.05–5.20$0.0188
verticalmedicalGPT-5.4 mini24/27 · 89%$0.0001518.210.9–25.5$0.00349
verticalmedicalGPT-5.6 Luna27/27 · 100%$0.000035876.163.3–90.1$0.00092
verticalmedicalGPT-5.6 Luna Pro27/27 · 100%$0.000309.187.23–11.2$0.00772
verticalmedicalGPT-5.6 Sol27/27 · 100%$0.001312.071.77–2.49$0.0336
verticalmedicalGPT-5.6 Terra27/27 · 100%$0.000309.097.43–11.2$0.00764
verticalmedicalgpt-oss 120B27/27 · 100%$0.0000293$0.0000209–$0.000039492.8$0.00076
verticalmedicalKimi K2.526/27 · 96%$0.001701.601.05–2.73$0.0395
verticalmedicalKimi K327/27 · 100%$0.001951.391.09–1.76$0.0534
verticalmedicalLlama 4 Maverick12/27 · 44%$0.000406.861.58–25.7$0.00345
verticalmedicalQwen3.5 Flash21/27 · 78%$0.000318.645.40–13.1$0.00667
verticalmedicalQwen3.8 Max27/27 · 100%$0.001152.371.98–2.75$0.0301