Published task benchmark

Security Alert Investigation Agent

Enrich an alert across identity, endpoint, network, and cloud tools and close or escalate it.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

IT and security operationsAnalysisEvaluated 9/11/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Security Alert Investigation Agent

Enrich an alert across identity, endpoint, network, and cloud tools and close or escalate it.

IT and security operationsPer alert
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
1,805 in · 929 out · 15.56s p50
$13.61 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
89%
Cheapest
Mistral Large 2407
123B · Off-the-shelf
1,933 in · 1,173 out · 30.18s p50
$288 /100K
Best current price · TrustedRouter
Compare 2 router rates · observed 9/11/2026
TrustedRouter · $0.5275 input · $1.58 output /1M
OpenRouter · $2.00 input · $6.00 output /1M
91%
GPT-5.4
Hosted · Off-the-shelf
1,514 in · 1,569 out · 16.49s p50
$2732 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
98%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small45%100/100$45.93
Qwen3 235B A22Bmid87%100/100$360
DeepSeek V3mid84%100/100$136
Mistral Large 2407mid91%100/100$288
Arcee Trinity Large Thinkingmid59%100/100$176
GPT-5.4frontier98%100/100$2732
Claude Opus 4.7frontier53%100/100$6133
Gemini 3.1 Pro Previewfrontier38%100/100$2550
Gemma 4 E4B ITsmall89%100/100$13.61
Granite 4.1 8Bsmall70%100/100$17.14
Ministral 8B Instruct 2410small78%100/100$45.39
Qwen3 4B Instruct 2507small28%100/100$147

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-11T07:22:02.309Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task