Benchmark · live

Model reliability by task.

Aggregate cost and reliability for the models teams actually run on Daslab, grouped by the kind of work — not who ran what or how much, just how well each model does the job and what it costs.

Provisional · aggregates from live runs · fixed task set · thin cells hidden
The frontier · cheapest reliable model per task

The most reliable spend per task: the cheapest model that still finishes 90% or more of its runs. Finishing means it ran without error; correctness graders are coming.

TaskFrontier pickReliabilityCost / task
Finance & ERPclaude-sonnet-5100%$0.160
Codingglm-5.2100%$0.185
Research & webclaude-haiku-4-5-20251001100%$0.006
Comms & CRMglm-5.296%$0.050
Ops & infraqwen3.7-max100%$0.198
Robotics & devicesgemini-3.5-flash100%$0.834
Data & sheetsqwen3.7-max92%$0.121
Finance & ERP
ModelReliabilityCost / task
gemini-3.5-flash68%$0.575
claude-sonnet-4-681%$0.320
claude-sonnet-5 100%$0.160
qwen3.5-35b-a3b °80%$0.004
Coding
ModelReliabilityCost / task
kimi-k3 °89%$2.999
claude-opus-4-895%$1.605
qwen3.7-max85%$1.583
gpt-5.592%$0.950
claude-sonnet-4-680%$0.660
claude-sonnet-5 °75%$0.468
kimi-k2.7-code73%$0.299
glm-5.2 °100%$0.185
gemini-3.5-flash 87%$0.064
qwen3.5-35b-a3b 69%$0.004
Research & web
ModelReliabilityCost / task
claude-opus-4-895%$0.851
kimi-k397%$0.580
gemini-3.5-flash65%$0.496
gpt-5.594%$0.477
claude-sonnet-4-684%$0.404
qwen3.7-max97%$0.317
kimi-k2.7-code89%$0.198
claude-sonnet-598%$0.184
glm-5.298%$0.042
kimi-k2.6100%$0.038
claude-haiku-4-5-20251001 100%$0.006
gpt-5-nano 87%$0.001
Comms & CRM
ModelReliabilityCost / task
claude-sonnet-595%$0.315
kimi-k3 °86%$0.241
qwen3.7-max °83%$0.152
gemini-3.5-flash51%$0.127
claude-sonnet-4-686%$0.103
glm-5.2 96%$0.050
Ops & infra
ModelReliabilityCost / task
gemini-3.5-flash °100%$0.201
qwen3.7-max °100%$0.198
Robotics & devices
ModelReliabilityCost / task
gemini-3.5-flash °100%$0.834
Data & sheets
ModelReliabilityCost / task
claude-sonnet-4-6 °75%$0.949
gpt-5.5 °100%$0.892
claude-sonnet-5 °100%$0.184
qwen3.7-max 92%$0.121
gemini-3.5-flash73%$0.097
qwen3.5-35b-a3b 82%$0.007

How this is measured. Every run on Daslab over the last 60 days is placed into a fixed task set by the tools it called, then aggregated by model. Reliability = the share of those runs that finished without error (running runs excluded); it does not yet judge whether the answer was correct. Cost / task = median spend per finished run. marks the cost / reliability frontier: no other model is both cheaper and more reliable. ° marks a provisional cell with few runs. A model needs at least five runs in a task to appear. Aggregates only: no usage volumes, no counts, no per-customer data.