Benchmark · live
Model reliability by task.
Aggregate cost and reliability for the models teams actually run on Daslab, grouped by the kind of work — not who ran what or how much, just how well each model does the job and what it costs.
Provisional · aggregates from live runs · fixed task set · thin cells hidden
The frontier · cheapest reliable model per task
The most reliable spend per task: the cheapest model that still finishes 90% or more of its runs. Finishing means it ran without error; correctness graders are coming.
| Task | Frontier pick | Reliability | Cost / task |
| Finance & ERP | claude-sonnet-5 | 100% | $0.160 |
| Coding | glm-5.2 | 100% | $0.185 |
| Research & web | claude-haiku-4-5-20251001 | 100% | $0.006 |
| Comms & CRM | glm-5.2 | 96% | $0.050 |
| Ops & infra | qwen3.7-max | 100% | $0.198 |
| Robotics & devices | gemini-3.5-flash | 100% | $0.834 |
| Data & sheets | qwen3.7-max | 92% | $0.121 |
Finance & ERP
| Model | Reliability | Cost / task |
| gemini-3.5-flash | 68% | $0.575 |
| claude-sonnet-4-6 | 81% | $0.320 |
| claude-sonnet-5 ★ | 100% | $0.160 |
| qwen3.5-35b-a3b ★ ° | 80% | $0.004 |
Coding
| Model | Reliability | Cost / task |
| kimi-k3 ° | 89% | $2.999 |
| claude-opus-4-8 | 95% | $1.605 |
| qwen3.7-max | 85% | $1.583 |
| gpt-5.5 | 92% | $0.950 |
| claude-sonnet-4-6 | 80% | $0.660 |
| claude-sonnet-5 ° | 75% | $0.468 |
| kimi-k2.7-code | 73% | $0.299 |
| glm-5.2 ★ ° | 100% | $0.185 |
| gemini-3.5-flash ★ | 87% | $0.064 |
| qwen3.5-35b-a3b ★ | 69% | $0.004 |
Research & web
| Model | Reliability | Cost / task |
| claude-opus-4-8 | 95% | $0.851 |
| kimi-k3 | 97% | $0.580 |
| gemini-3.5-flash | 65% | $0.496 |
| gpt-5.5 | 94% | $0.477 |
| claude-sonnet-4-6 | 84% | $0.404 |
| qwen3.7-max | 97% | $0.317 |
| kimi-k2.7-code | 89% | $0.198 |
| claude-sonnet-5 | 98% | $0.184 |
| glm-5.2 | 98% | $0.042 |
| kimi-k2.6 | 100% | $0.038 |
| claude-haiku-4-5-20251001 ★ | 100% | $0.006 |
| gpt-5-nano ★ | 87% | $0.001 |
Comms & CRM
| Model | Reliability | Cost / task |
| claude-sonnet-5 | 95% | $0.315 |
| kimi-k3 ° | 86% | $0.241 |
| qwen3.7-max ° | 83% | $0.152 |
| gemini-3.5-flash | 51% | $0.127 |
| claude-sonnet-4-6 | 86% | $0.103 |
| glm-5.2 ★ | 96% | $0.050 |
Ops & infra
| Model | Reliability | Cost / task |
| gemini-3.5-flash ° | 100% | $0.201 |
| qwen3.7-max ★ ° | 100% | $0.198 |
Robotics & devices
| Model | Reliability | Cost / task |
| gemini-3.5-flash ★ ° | 100% | $0.834 |
Data & sheets
| Model | Reliability | Cost / task |
| claude-sonnet-4-6 ° | 75% | $0.949 |
| gpt-5.5 ° | 100% | $0.892 |
| claude-sonnet-5 ★ ° | 100% | $0.184 |
| qwen3.7-max ★ | 92% | $0.121 |
| gemini-3.5-flash | 73% | $0.097 |
| qwen3.5-35b-a3b ★ | 82% | $0.007 |