Published task benchmark

Role Transfer Agent

Apply role changes across HRIS, permissions, equipment, training, and reporting relationships.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

People and HR operationsAnalysisEvaluated 9/11/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Role Transfer Agent

Apply role changes across HRIS, permissions, equipment, training, and reporting relationships.

People and HR operationsPer transfer
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
902 in · 1,559 out · 25.00s p50
$18.35 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
83%
Best accuracy
Mistral Large 2407
123B · Off-the-shelf
1,000 in · 1,336 out · 34.53s p50
$264 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.5275 input · $1.58 output /1M
OpenRouter · $2.00 input · $6.00 output /1M
78%
GPT-5.4
Hosted · Off-the-shelf
813 in · 1,943 out · 17.36s p50
$3118 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
44%
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small15%100/100$48.43
Qwen3 235B A22Bmid39%100/100$353
DeepSeek V3mid70%100/100$148
Mistral Large 2407mid78%100/100$264
Arcee Trinity Large Thinkingmid32%100/100$164
GPT-5.4frontier44%100/100$3118
Claude Opus 4.7frontier11%100/100$5694
Gemini 3.1 Pro Previewfrontier9%100/100$2567
Gemma 4 E4B ITsmall83%100/100$18.35
Granite 4.1 8Bsmall48%100/100$14.23
Ministral 8B Instruct 2410small54%100/100$32.17
Qwen3 4B Instruct 2507small6%100/100$136

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T10:06:02.086Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task