Published task benchmark

Vendor Entity Deduplication

Given vendor-master records, identify likely duplicates across name, address, tax ID, and payment details while distinguishing related but separate legal entities.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Finance & AccountingAnalysisEvaluated 9/10/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Vendor Entity Deduplication

Given vendor-master records, identify likely duplicates across name, address, tax ID, and payment details while distinguishing related but separate legal entities.

Finance & AccountingPer vendor record
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
1,175 in · 1,512 out · 24.32s p50
$18.43 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
71%
Cheapest
DeepSeek V3
Hosted · Off-the-shelf
1,022 in · 1,176 out · 9.50s p50
$145 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.3376 input · $0.9390 output /1M
OpenRouter · $0.2900 input · $1.14 output /1M
73%
GPT-5.4
Hosted · Off-the-shelf
984 in · 1,892 out · 15.44s p50
$3084 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
77%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small19%100/100$51.71
Qwen3 235B A22Bmid56%100/100$398
DeepSeek V3mid73%100/100$145
Mistral Large 2407mid53%100/100$337
Arcee Trinity Large Thinkingmid57%100/100$169
GPT-5.4frontier77%100/100$3084
Claude Opus 4.7frontier34%100/100$5787
Gemini 3.1 Pro Previewfrontier1%99/100$2627
Gemma 4 E4B ITsmall71%100/100$18.43
Granite 4.1 8Bsmall36%100/100$19.47
Ministral 8B Instruct 2410small11%100/100$43.45
Qwen3 4B Instruct 2507small29%100/100$140

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T08:59:56.429Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task