Published task benchmark

Prior Authorization Evidence Assembly

Given a coverage policy and synthetic chart, map documented evidence to each requirement and identify unmet or missing documentation without deciding medical necessity.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Healthcare & Life SciencesAnalysisEvaluated 9/10/202697 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Prior Authorization Evidence Assembly

Given a coverage policy and synthetic chart, map documented evidence to each requirement and identify unmet or missing documentation without deciding medical necessity.

Healthcare & Life SciencesPer chart
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
1,015 in · 1,208 out · 18.66s p50
$14.89 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
42%
Cheapest
DeepSeek V3
Hosted · Off-the-shelf
961 in · 1,104 out · 10.25s p50
$136 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.3376 input · $0.9390 output /1M
OpenRouter · $0.2900 input · $1.14 output /1M
65%
GPT-5.4
Hosted · Off-the-shelf
946 in · 1,944 out · 15.53s p50
$3153 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
56%
All 12 model results and methodology

Generated 97/100 examples — partial set accepted at the success threshold

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small26%97/97$47.11
Qwen3 235B A22Bmid10%97/97$323
DeepSeek V3mid65%97/97$136
Mistral Large 2407mid28%97/97$295
Arcee Trinity Large Thinkingmid21%97/97$172
GPT-5.4frontier56%97/97$3153
Claude Opus 4.7frontier52%97/97$5684
Gemini 3.1 Pro Previewfrontier2%94/97$2607
Gemma 4 E4B ITsmall42%97/97$14.89
Granite 4.1 8Bsmall16%97/97$15.74
Ministral 8B Instruct 2410small23%97/97$33.90
Qwen3 4B Instruct 2507small3%97/97$123

LLM-judge pass rate on 97 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T08:55:02.675Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task