Published task benchmark

Last-Mile Instruction Clarification

Given a customer's delivery note, rewrite it as concise driver steps while preserving access constraints and flagging ambiguous or unsafe instructions.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

E-commerce & LogisticsGenerationEvaluated 9/9/2026100 evaluation examplesFallback pricing checked September 11, 2026
Generation

Last-Mile Instruction Clarification

Given a customer's delivery note, rewrite it as concise driver steps while preserving access constraints and flagging ambiguous or unsafe instructions.

E-commerce & LogisticsPer delivery note
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Nemotron Nano 9B v2
9B · Off-the-shelf
239 in · 1,169 out · 6.82s p50
$28.32 /100K
AWS Bedrock temporary list price
View token rates
AWS Bedrock · $0.0600 input · $0.2300 output /1M
36%
DeepSeek V3
Hosted · Off-the-shelf
233 in · 121 out · 1.99s p50
$19.23 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.3376 input · $0.9390 output /1M
OpenRouter · $0.2900 input · $1.14 output /1M
69%
GPT-5.4
Hosted · Off-the-shelf
223 in · 128 out · 2.22s p50
$248 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
94%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small36%100/100$28.32
Qwen3 235B A22Bmid62%100/100$39.99
DeepSeek V3mid69%100/100$19.23
Mistral Large 2407mid66%100/100$43.57
Arcee Trinity Large Thinkingmid67%100/100$61.20
GPT-5.4frontier94%100/100$248
Claude Opus 4.7frontier94%100/100$1107
Gemini 3.1 Pro Previewfrontier82%100/100$1305
Gemma 4 E4B ITsmall21%100/100$1.76
Granite 4.1 8Bsmall23%100/100$3.71
Ministral 8B Instruct 2410small15%100/100$5.70
Qwen3 4B Instruct 2507small27%100/100$16.24

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-09T11:07:02.682Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task