Published task benchmark

GDPR DPA Article 28 Validation

Given a data-processing addendum and an Article 28 requirements checklist, classify each requirement as present, missing, or ambiguous with supporting text spans.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Legal & ComplianceClassificationEvaluated 9/10/2026100 evaluation examplesFallback pricing checked September 11, 2026
Classification

GDPR DPA Article 28 Validation

Given a data-processing addendum and an Article 28 requirements checklist, classify each requirement as present, missing, or ambiguous with supporting text spans.

Legal & CompliancePer addendum
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Nemotron Nano 9B v2
9B · Off-the-shelf
710 in · 1,035 out · 6.33s p50
$28.07 /100K
AWS Bedrock temporary list price
View token rates
AWS Bedrock · $0.0600 input · $0.2300 output /1M
24%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
694 in · 362 out · 4.54s p50
$97.46 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.4550 input · $1.82 output /1M
TrustedRouter · $0.6330 input · $2.11 output /1M
26%
Claude Opus 4.7
Hosted · Off-the-shelf
1,132 in · 535 out · 6.17s p50
$1904 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $5.00 input · $25.00 output /1M
TrustedRouter · $5.28 input · $26.38 output /1M
84%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small24%100/100$28.07
Qwen3 235B A22Bmid26%100/100$97.46
DeepSeek V3mid19%100/100$53.30
Mistral Large 2407mid19%100/100$125
Arcee Trinity Large Thinkingmid24%100/100$138
GPT-5.4frontier70%100/100$1112
Claude Opus 4.7frontier84%100/100$1904
Gemini 3.1 Pro Previewfrontier43%100/100$2167
Gemma 4 E4B ITsmall14%100/100$3.66
Granite 4.1 8Bsmall19%100/100$7.41
Ministral 8B Instruct 2410small12%100/100$18.69
Qwen3 4B Instruct 2507small20%100/100$53.55

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T09:02:02.487Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task