← All models

Llama 4 Maverick

groq/llama-4-maverick

Llama 4 Maverick is the larger Llama 4 on Groq. It is suited for quality-sensitive chat that still needs interactive latency.

Pricing & limits

Provider list price, passed through untouched. The only SwitchGate fee is 4.9% when you top up credits.

ProviderGroq
Input price$0.2 / 1M tokens
Output price$0.6 / 1M tokens
Context window131K tokens
Added to catalog2026-04-10
Tokens routed (30d)97.5B (placeholder)
Cache hits0 credits, always
  • Same governance as every model: per-user budgets, model locks, burnable keys and the kill switch apply automatically.
  • Cross-provider failover: route around outages to sibling models with one retry policy.
  • Exact cost per call on one ledger, whoever on your team calls it.

Call it in one line

Python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.switchgate.ai/v1",
    api_key="sg-live-…",
)

r = client.chat.completions.create(
    model="groq/llama-4-maverick",
    messages=[{"role": "user", "content": "Bonjour!"}],
)

Works with any OpenAI-compatible SDK — JavaScript, Go, cURL and more on the quickstart. Compare quality and cost on benchmarks and rankings.

More from Groq

Route Llama 4 Maverick in the next five minutes

Free trial credits, list-price pass-through, and a ledger that shows exactly what each call cost.