Published task benchmark

Product Attribute Tagging

Given a product title, description, and allowed taxonomy, assign supported color, material, style, occasion, and feature tags without inventing attributes.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

E-commerce & LogisticsExtractionEvaluated 9/9/2026100 evaluation examplesFallback pricing checked September 11, 2026
Extraction

Product Attribute Tagging

Given a product title, description, and allowed taxonomy, assign supported color, material, style, occasion, and feature tags without inventing attributes.

E-commerce & LogisticsPer product
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Qwen3 4B Instruct 2507
4B · Off-the-shelf
315 in · 53 out · 2.16s p50
$8.14 /100K
AWS Bedrock temporary proxy price
View token rates
AWS Bedrock · $0.1545 input · $0.6180 output /1M
46%
Mistral Large 2407
123B · Off-the-shelf
349 in · 84 out · 2.16s p50
$31.70 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.5275 input · $1.58 output /1M
OpenRouter · $2.00 input · $6.00 output /1M
59%
GPT-5.4
Hosted · Off-the-shelf
303 in · 62 out · 1.25s p50
$169 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
74%
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small44%100/100$17.38
Qwen3 235B A22Bmid36%100/100$24.34
DeepSeek V3mid57%100/100$15.01
Mistral Large 2407mid59%100/100$31.70
Arcee Trinity Large Thinkingmid56%100/100$88.33
GPT-5.4frontier74%100/100$169
Claude Opus 4.7frontier62%100/100$478
Gemini 3.1 Pro Previewfrontier61%100/100$958
Gemma 4 E4B ITsmall36%100/100$1.67
Granite 4.1 8Bsmall40%100/100$2.06
Ministral 8B Instruct 2410small33%100/100$5.73
Qwen3 4B Instruct 2507small46%100/100$8.14

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-09T09:53:02.599Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task