Published task benchmark

Appointment Management Agent

Find feasible slots and book, reschedule, or cancel an appointment under policy constraints.

Compare measured quality, workload cost, and latency across model tiers. Results are task-specific and directional—not a general model ranking.

Customer service and account operationsAnalysisEvaluated 9/10/2026100 evaluation examplesFallback pricing checked September 11, 2026
Analysis

Appointment Management Agent

Find feasible slots and book, reschedule, or cancel an appointment under policy constraints.

Customer service and account operationsPer appointment
Not benchmarked
No result on this tier
Token profile pending
/100K
Run this tier to add it
Gemma 4 E4B IT
E4B · Off-the-shelf
685 in · 910 out · 14.85s p50
$11.05 /100K
Current price · TrustedRouter
View token rates · observed 9/11/2026
TrustedRouter · $0.0211 input · $0.1055 output /1M
25%
Qwen3 235B A22B
235B (22B active) · Off-the-shelf
652 in · 1,226 out · 18.13s p50
$253 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $0.4550 input · $1.82 output /1M
TrustedRouter · $0.6330 input · $2.11 output /1M
37%
GPT-5.4
Hosted · Off-the-shelf
538 in · 631 out · 5.94s p50
$1081 /100K
Best current price · OpenRouter
Compare 2 router rates · observed 9/11/2026
OpenRouter · $2.50 input · $15.00 output /1M
TrustedRouter · $2.64 input · $15.82 output /1M
76%
Best accuracy
All 12 model results and methodology

Generated 100/100 examples

ModelTierQualityJudgedScenario cost
Nemotron Nano 9B v2small24%100/100$38.18
Qwen3 235B A22Bmid37%100/100$253
DeepSeek V3mid36%100/100$89.03
Mistral Large 2407mid32%100/100$195
Arcee Trinity Large Thinkingmid19%100/100$155
GPT-5.4frontier76%100/100$1081
Claude Opus 4.7frontier75%100/100$3449
Gemini 3.1 Pro Previewfrontier60%100/100$1746
Gemma 4 E4B ITsmall25%100/100$11.05
Granite 4.1 8Bsmall9%100/100$9.32
Ministral 8B Instruct 2410small13%100/100$23.57
Qwen3 4B Instruct 2507small14%100/100$103

LLM-judge pass rate on 100 synthetic examples. Generator: gpt-5.2. Judge: gpt-5.2. Evaluated 2026-09-10T19:45:02.279Z.

Directional: measured on a synthetic eval set generated by drydock. Cost/latency are not yet captured for taskrouter-run benchmarks.

drydock

Have a neighboring workload?

Describe it in Task Explorer. If it is not measured yet, you can request a benchmark and follow future results.

Explore another task