Published task benchmark

Duplicate Engineering Ticket Detection

Given a new issue and an existing issue corpus, rank likely duplicates and distinguish same symptom, same root cause, and merely related requests.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Engineering & ProductAnalysisEvaluated 9/10/202692 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Duplicate Engineering Ticket Detection

Given a new issue and an existing issue corpus, rank likely duplicates and distinguish same symptom, same root cause, and merely related requests.

Engineering & ProductPer ticket
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Granite 4.1 8B
8B · Off-the-shelf
818 in · 1,144 out · 17.37s p50
$15.53 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/4/2026
OpenRouter · $0.0500 input · $0.1000 output /1M
TrustedRouter · $0.0527 input · $0.1055 output /1M
88%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
867 in · 1,492 out · 20.29s p50
$311 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.4550 input · $1.82 output /1M
TrustedRouter · $0.6330 input · $2.11 output /1M
90%
GPT-5.4
Hosted · Off-the-shelf
816 in · 1,345 out · 12.01s p50
$2222 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
99%
Best accuracy
All 12 model results and methodology

Generated 92/100 examples — partial set accepted at the success threshold

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small54%92/92$48.37
Qwen3 235B A22Bmid90%92/92$311
DeepSeek V3mid89%92/92$125
Mistral Large 2407mid76%92/92$231
Arcee Trinity Large Thinkingmid60%92/92$149
GPT-5.4frontier99%92/92$2222
Claude Opus 4.7frontier95%92/92$5038
Gemini 3.1 Pro Previewfrontier8%92/92$2567
Gemma 4 E4B ITsmall85%92/92$14.23
Granite 4.1 8Bsmall88%92/92$15.53
Ministral 8B Instruct 2410small34%92/92$28.47
Qwen3 4B Instruct 2507small77%92/92$114

LLM-judge pass rate on 92 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T08:55:02.630Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task