Published task benchmark

Project Status Sync Agent

Reconcile status across tickets, documents, and messages and update the project system with evidence.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Office and knowledge-work automationAnalysisEvaluated 9/11/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Project Status Sync Agent

Reconcile status across tickets, documents, and messages and update the project system with evidence.

Office and knowledge-work automationPer project
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
1,202 in · 1,056 out · 16.83s p50
$13.68 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
32%
Cheapest
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
1,165 in · 1,321 out · 11.81s p50
$293 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.4550 input · $1.82 output /1M
TrustedRouter · $0.6330 input · $2.11 output /1M
47%
GPT-5.4
Hosted · Off-the-shelf
1,046 in · 1,389 out · 10.92s p50
$2345 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
63%
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small12%100/100$48.01
Qwen3 235B A22Bmid47%100/100$293
DeepSeek V3mid43%100/100$118
Mistral Large 2407mid41%100/100$237
Arcee Trinity Large Thinkingmid20%100/100$179
GPT-5.4frontier63%100/100$2345
Claude Opus 4.7frontier60%100/100$5556
Gemini 3.1 Pro Previewfrontier2%99/100$2633
Gemma 4 E4B ITsmall32%100/100$13.68
Granite 4.1 8Bsmall17%100/100$15.71
Ministral 8B Instruct 2410small12%100/100$33.13
Qwen3 4B Instruct 2507small19%100/100$118

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T12:32:02.434Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task