Published task benchmark

Email Unsubscribe Reason Classification

Given unsubscribe comments, assign one supplied reason label per response and use the fallback for ambiguous or irrelevant feedback.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Sales & MarketingClassificationEvaluated 9/9/2026100 evaluation examplesFallback pricing checked September 11, 2026
Classification

Email Unsubscribe Reason Classification

Given unsubscribe comments, assign one supplied reason label per response and use the fallback for ambiguous or irrelevant feedback.

Sales & MarketingPer feedback
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
199 in · 22 out · 0.39s p50
$0.6520 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
99%
Cheapest
Arcee Trinity Large Thinking
400B · Off-the-shelf
190 in · 378 out · 0.87s p50
$34.99 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.2500 input · $0.8000 output /1M
TrustedRouter · $0.2638 input · $0.8440 output /1M
99%
Claude Opus 4.7
Hosted · Off-the-shelf
336 in · 58 out · 1.73s p50
$313 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $5.00 input · $25.00 output /1M
TrustedRouter · $5.28 input · $26.38 output /1M
100%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small98%100/100$7.10
Qwen3 235B A22Bmid96%100/100$15.88
DeepSeek V3mid93%100/100$7.99
Mistral Large 2407mid97%100/100$16.72
Arcee Trinity Large Thinkingmid99%100/100$34.99
GPT-5.4frontier99%100/100$94.25
Claude Opus 4.7frontier100%100/100$313
Gemini 3.1 Pro Previewfrontier100%100/100$564
Gemma 4 E4B ITsmall99%100/100$0.6520
Granite 4.1 8Bsmall95%100/100$1.30
Ministral 8B Instruct 2410small91%100/100$3.42
Qwen3 4B Instruct 2507small94%100/100$5.52

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-09T02:47:02.690Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task