Expense Policy Decision Explanation
Given an expense claim and the applicable policy, classify it as approve, reject, or review and return the exact policy rule supporting the decision.
View token rates
Compare 2 router rates · observed 9/11/2026
Compare 2 router rates · observed 9/11/2026
All 12 model results and methodology
Generated 100/100 examples
| Model | Tier | Quality | Judged | Scenario cost |
|---|---|---|---|---|
| Nemotron Nano 9B v2 | small | 18% | 100/100 | $13.44 |
| Qwen3 235B A22B | mid | 24% | 100/100 | $25.21 |
| DeepSeek V3 | mid | 40% | 100/100 | $12.46 |
| Mistral Large 2407 | mid | 26% | 100/100 | $25.74 |
| Arcee Trinity Large Thinking | mid | 31% | 100/100 | $105 |
| GPT-5.4 | frontier | 83% | 100/100 | $189 |
| Claude Opus 4.7 | frontier | 31% | 100/100 | $420 |
| Gemini 3.1 Pro Preview | frontier | 55% | 100/100 | $807 |
| Gemma 4 E4B IT | small | 23% | 100/100 | $1.04 |
| Granite 4.1 8B | small | 28% | 100/100 | $2.07 |
| Ministral 8B Instruct 2410 | small | 25% | 100/100 | $5.68 |
| Qwen3 4B Instruct 2507 | small | 31% | 100/100 | $9.98 |
LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-08T22:06:03.517Z.
Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.
drydock