Published task benchmark

Grounded Contract Question Answering

Answer questions only from supplied authoritative contract chunks, cite the supporting chunk, and abstain when the answer is absent or superseded. Available model: neurometric/grounded-document-qa.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Legal & ComplianceExtractionEvaluated 9/8/2026100 evaluation examplesFallback pricing checked September 11, 2026
Extraction

Grounded Contract Question Answering

Answer questions only from supplied authoritative contract chunks, cite the supporting chunk, and abstain when the answer is absent or superseded. Available model: neurometric/grounded-document-qa.

Legal & CompliancePer contract
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Granite 4.1 8B
8B · Off-the-shelf
406 in · 180 out · 2.19s p50
$3.83 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/4/2026
OpenRouter · $0.0500 input · $0.1000 output /1M
TrustedRouter · $0.0527 input · $0.1055 output /1M
27%
Mistral Large 2407
123B · Off-the-shelf
459 in · 200 out · 5.84s p50
$55.86 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.5275 input · $1.58 output /1M
OpenRouter · $2.00 input · $6.00 output /1M
63%
GPT-5.4
Hosted · Off-the-shelf
400 in · 164 out · 2.13s p50
$346 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
84%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small13%100/100$13.96
Qwen3 235B A22Bmid59%100/100$48.41
DeepSeek V3mid53%100/100$19.58
Mistral Large 2407mid63%100/100$55.86
Arcee Trinity Large Thinkingmid39%100/100$109
GPT-5.4frontier84%100/100$346
Claude Opus 4.7frontier82%100/100$1582
Gemini 3.1 Pro Previewfrontier59%100/100$798
Gemma 4 E4B ITsmall17%100/100$1.79
Granite 4.1 8Bsmall27%100/100$3.83
Ministral 8B Instruct 2410small14%100/100$7.72
Qwen3 4B Instruct 2507small14%100/100$10.07

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-08T20:45:02.796Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task