Benchmarks
Independent, reproducible measurements of the knobs you can actually set on a SwitchGate request: models, providers and routing budgets. Every score links to the configuration, costs and telemetry behind it.
6 benchmarks · 2,464,375 task evaluations · placeholder data · last run Aug 12, 2026
Agents & tools
GateBench Airline →
Multi-turn service agents making tool calls under strict policy constraints.
110 models · last run Aug 11, 2026
Quality
81.3%
claude-opus-4-1
Value
$0.015
gemini-3-flash
Speed
1.8m
nova-lite
GateBench Retail →
Long-horizon shopping agents scored on order accuracy and refund policy compliance.
96 models · last run Aug 11, 2026
Quality
74.6%
gpt-5.2
Value
$0.011
mistral-small-4
Speed
2.1m
llama-4-scout
Reasoning
GPQA Diamond →
Graduate-level science questions that resist retrieval and reward careful reasoning.
110 models · last run Aug 11, 2026
Quality
94.3%
gemini-3-pro
Value
$0.021
gpt-5.2
Speed
41s
gpt-5-mini
MathArena 2026 →
Competition math scored with exact-answer verification, no partial credit.
84 models · last run Aug 11, 2026
Quality
88.1%
gpt-5.2
Value
$0.017
deepseek-r1
Speed
56s
grok-4-fast
Search
BrowseComp →
Hard-to-locate facts on the live web, scored on persistent multi-step research.
4 models · last run Aug 11, 2026
Quality
89.0%
claude-opus-4-1
Value
$0.99
claude-opus-4-1
Speed
1.9m
claude-opus-4-1
DeepSearchQA →
Questions whose answers are lists, scored for exhaustive retrieval with no padding.
4 models · last run Aug 11, 2026
Quality
76.5%
claude-opus-4-1
Value
$0.091
deepseek-v3.2
Speed
1.6m
gpt-5.2
Placeholder scores for design purposes; reproducible harnesses publish with the live catalog.
Benchmark it on your own traffic
Every request through the gateway is traced with exact cost and latency, so your production numbers become your benchmark.