Classification
Phishing Triage Agent
Inspect reported messages and telemetry, classify risk, contain artifacts, and notify affected users.
IT and security operationsPer message
Not benchmarked
No result on this tier
Token profile pending
— /100K
Run this tier to add it
—
Nemotron Nano 9B v2
9B · Off-the-shelf
679 in · 537 out · 4.04s p50
$16.43 /100K
AWS Bedrock temporary list price
View token rates
AWS Bedrock · $0.0600 input · $0.2300 output /1M
78%
Best accuracy
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
647 in · 48 out · 0.71s p50
$38.17 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.4550 input · $1.82 output /1M
TrustedRouter · $0.6330 input · $2.11 output /1M
76%
GPT-5.4
Hosted · Off-the-shelf
600 in · 49 out · 1.55s p50
$224 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
75%
All 12 model results and methodology
Generated 100/100 examples
| Model | Tier | Quality | Judged | Scenario cost |
|---|---|---|---|---|
| Nemotron Nano 9B v2 | small | 78% | 100/100 | $16.43 |
| Qwen3 235B A22B | mid | 76% | 100/100 | $38.17 |
| DeepSeek V3 | mid | 66% | 100/100 | $18.23 |
| Mistral Large 2407 | mid | 75% | 100/100 | $48.06 |
| Arcee Trinity Large Thinking | mid | 59% | 100/100 | $88.94 |
| GPT-5.4 | frontier | 75% | 100/100 | $224 |
| Claude Opus 4.7 | frontier | 57% | 100/100 | $812 |
| Gemini 3.1 Pro Preview | frontier | 43% | 100/100 | $787 |
| Gemma 4 E4B IT | small | 70% | 100/100 | $1.83 |
| Granite 4.1 8B | small | 76% | 100/100 | $3.53 |
| Ministral 8B Instruct 2410 | small | 67% | 100/100 | $11.03 |
| Qwen3 4B Instruct 2507 | small | 74% | 100/100 | $15.81 |
LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T05:58:03.763Z.
Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.
drydock