Published task benchmark

Month-End Close Agent

Track dependencies, gather evidence, run close checks, and advance or block checklist items.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Finance, accounting, and procurement operationsAnalysisEvaluated 9/11/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Month-End Close Agent

Track dependencies, gather evidence, run close checks, and advance or block checklist items.

Finance, accounting, and procurement operationsPer checklist
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
1,514 in · 1,365 out · 22.34s p50
$17.60 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
48%
Cheapest
DeepSeek V3
Hosted · Off-the-shelf
1,354 in · 1,225 out · 14.91s p50
$156 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.2500 input · $1.00 output /1M
TrustedRouter · $0.3376 input · $0.9390 output /1M
63%
Claude Opus 4.7
Hosted · Off-the-shelf
2,031 in · 1,923 out · 19.89s p50
$5823 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $5.00 input · $25.00 output /1M
TrustedRouter · $5.28 input · $26.38 output /1M
56%
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small38%100/100$50.79
Qwen3 235B A22Bmid56%100/100$378
DeepSeek V3mid63%100/100$156
Mistral Large 2407mid37%100/100$350
Arcee Trinity Large Thinkingmid14%100/100$189
GPT-5.4frontier42%100/100$3281
Claude Opus 4.7frontier56%100/100$5823
Gemini 3.1 Pro Previewfrontier6%100/100$2687
Gemma 4 E4B ITsmall48%100/100$17.60
Granite 4.1 8Bsmall36%100/100$18.55
Ministral 8B Instruct 2410small13%100/100$43.56
Qwen3 4B Instruct 2507small26%100/100$140

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T00:50:03.408Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task