Published task benchmark

Payroll Exception Agent

Investigate missing or incorrect pay using time, rate, deduction, and payroll records.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

People and HR operationsAnalysisEvaluated 9/11/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Payroll Exception Agent

Investigate missing or incorrect pay using time, rate, deduction, and payroll records.

People and HR operationsPer paycheck
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
882 in · 1,155 out · 18.07s p50
$14.05 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
41%
DeepSeek V3
Hosted · Off-the-shelf
784 in · 954 out · 11.73s p50
$115 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.2500 input · $1.00 output /1M
TrustedRouter · $0.3376 input · $0.9390 output /1M
38%
GPT-5.4
Hosted · Off-the-shelf
750 in · 1,132 out · 10.61s p50
$1886 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
82%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small35%100/100$42.73
Qwen3 235B A22Bmid28%100/100$318
DeepSeek V3mid38%100/100$115
Mistral Large 2407mid28%100/100$219
Arcee Trinity Large Thinkingmid19%100/100$165
GPT-5.4frontier82%100/100$1886
Claude Opus 4.7frontier72%100/100$4303
Gemini 3.1 Pro Previewfrontier41%100/100$2338
Gemma 4 E4B ITsmall41%100/100$14.05
Granite 4.1 8Bsmall11%100/100$13.87
Ministral 8B Instruct 2410small17%100/100$28.59
Qwen3 4B Instruct 2507small5%100/100$123

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T04:46:02.343Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task