Published task benchmark

Employee Engagement Survey Theming

Given anonymized survey comments, assign supplied themes, quantify prevalence, and extract representative evidence without reidentifying respondents.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

HR & Internal OpsAnalysisEvaluated 9/10/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Employee Engagement Survey Theming

Given anonymized survey comments, assign supplied themes, quantify prevalence, and extract representative evidence without reidentifying respondents.

HR & Internal OpsPer survey
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Nemotron Nano 9B v2
9B · Off-the-shelf
682 in · 1,866 out · 12.16s p50
$47.01 /100K
AWS Bedrock temporary list price
View token rates
AWS Bedrock · $0.0600 input · $0.2300 output /1M
37%
Arcee Trinity Large Thinking
400B · Off-the-shelf
650 in · 1,762 out · 5.50s p50
$157 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.2500 input · $0.8000 output /1M
TrustedRouter · $0.2638 input · $0.8440 output /1M
27%
GPT-5.4
Hosted · Off-the-shelf
652 in · 1,870 out · 17.19s p50
$2968 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
62%
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small37%100/100$47.01
Qwen3 235B A22Bmid11%100/100$317
DeepSeek V3mid26%100/100$128
Mistral Large 2407mid24%100/100$221
Arcee Trinity Large Thinkingmid27%100/100$157
GPT-5.4frontier62%100/100$2968
Claude Opus 4.7frontier39%100/100$5477
Gemini 3.1 Pro Previewfrontier1%100/100$2533
Gemma 4 E4B ITsmall26%100/100$14.72
Granite 4.1 8Bsmall14%100/100$15.93
Ministral 8B Instruct 2410small6%100/100$24.00
Qwen3 4B Instruct 2507small5%100/100$114

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T08:50:02.877Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task