Benchmark · live

Model reliability by task

Aggregate cost and reliability for the models teams actually run on Daslab, grouped by the kind of work — not who ran what or how much, just how well each model does the job and what it costs.

Provisional · aggregates from live runs · fixed task set · thin cells hidden
The frontier · cheapest reliable model per task

The most reliable spend per task: the cheapest model that still finishes 90% or more of its runs. Finishing means it ran without error; correctness graders are coming.

TaskFrontier pickReliabilityCost / task
Finance & ERPgemini-3.7-flash100%$0.053
Codinggemini-3.7-flash100%$0.276
Research & webdeepseek-v4-flash-0731100%$0.003
Comms & CRMgemini-3.7-flash100%$0.035
Ops & infragemini-3.7-flash90–100%$0.030
Finance & ERP
ModelReliabilityTool errorsCost / task
kimi-k3 °70–80%0–10%$3.354
gemini-3.5-flash60–70%0–10%$0.413
claude-sonnet-590–100%0–10%$0.413
gemini-3.7-flash 100%$0.053
Coding
ModelReliabilityTool errorsCost / task
kimi-k3 °90–100%0–10%$0.621
gemini-3.7-flash °100%$0.276
Research & web
ModelReliabilityTool errorsCost / task
kimi-k390–100%0–10%$0.559
gemini-3.5-flash40–50%0–10%$0.438
claude-sonnet-590–100%0–10%$0.189
qwen3.7-max °100%0–10%$0.129
kimi-k2.7-code80–90%$0.109
auto90–100%$0.040
gemini-3.7-flash90–100%$0.039
glm-5.290–100%0–10%$0.028
deepseek-v4-flash-0731 100%0–10%$0.003
Comms & CRM
ModelReliabilityTool errorsCost / task
claude-sonnet-5 °90–100%0–10%$0.289
gemini-3.5-flash50–60%0–10%$0.069
glm-5.2 °90–100%0–10%$0.063
gemini-3.7-flash 100%$0.035
Ops & infra
ModelReliabilityTool errorsCost / task
gemini-3.5-flash90–100%0–10%$0.242
gemini-3.7-flash °90–100%$0.030

How this is measured. Every run on Daslab over the last 60 days is placed into a fixed task set by the tools it called, then aggregated by the model its turns actually ran on — a run that switched models mid-way is evidence about routing rather than about any one model, so it is excluded rather than credited to one. Reliability = the share of runs that finished without error (running runs excluded); it does not yet judge whether the answer was correct. Tool errors = the share of a model's tool calls that the tool rejected. Cost / task = median spend per finished run. marks the cost / reliability frontier: no other model is both cheaper and more reliable. ° marks a provisional cell.

Why the numbers are banded. Percentages are published in ten-point ranges on purpose. An exact percentage can be inverted — only certain values are reachable from a given number of runs, so “80%” would quietly disclose the sample it came from. Bands collapse many fractions onto one label. We publish how well the models do, never how much we run: no volumes, no counts, no per-customer data. Metrics that would need to know what the user wanted — how many turns a task took, whether depth was wanted — are deliberately absent, because fewer turns is not obviously better and a leaderboard that says so would reward models that give up early.