Llama 4 Scout
groq/llama-4-scout
Llama 4 Scout on Groq hardware is the speed play. It is suited for real-time voice, autocomplete and agent steps where 800+ tok/s changes the product.
Pricing & limits
Provider list price, passed through untouched. The only SwitchGate fee is 4.9% when you top up credits.
| Provider | Groq |
|---|---|
| Input price | $0.11 / 1M tokens |
| Output price | $0.34 / 1M tokens |
| Context window | 131K tokens |
| Added to catalog | 2026-04-10 |
| Tokens routed (30d) | 212.8B (placeholder) |
| Cache hits | 0 credits, always |
- Same governance as every model: per-user budgets, model locks, burnable keys and the kill switch apply automatically.
- Cross-provider failover: route around outages to sibling models with one retry policy.
- Exact cost per call on one ledger, whoever on your team calls it.
Call it in one line
from openai import OpenAI
client = OpenAI(
base_url="https://api.switchgate.ai/v1",
api_key="sg-live-…",
)
r = client.chat.completions.create(
model="groq/llama-4-scout",
messages=[{"role": "user", "content": "Bonjour!"}],
)
Works with any OpenAI-compatible SDK — JavaScript, Go, cURL and more on the quickstart. Compare quality and cost on benchmarks and rankings.
More from Groq
Route Llama 4 Scout in the next five minutes
Free trial credits, list-price pass-through, and a ledger that shows exactly what each call cost.