Groq's LPU silicon delivers some of the fastest token streams in inference, but on a deliberately narrow model list — and two of its best-known Llama tiers are enterprise-quoted with no public price. If you are shopping for an alternative, the real question is which constraint you are escaping: model coverage, per-token price, or rate limits. This page compares 8 credible alternatives with prices dated 2026-09-22.
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Key facts
- Groq's public self-serve list (as of 2026-09-22):
openai/gpt-oss-120bat $0.15 / $0.60 per million tokens at a listed ~500 tokens/s,openai/gpt-oss-20bat $0.075 / $0.30 at ~1,000 tokens/s, andqwen/qwen3.8-27bat $0.80 / $4.00 — per Groq's supported-models page. - The Llama gap: Groq lists Llama 3.1 8B Instant and Llama 3.3 70B Versatile as enterprise-only ("Contact Sales") with no public per-token rate, while SambaNova lists Llama 3.3 70B at $0.60 / $1.20 (SambaNova pricing).
- Speed rival: Cerebras lists gpt-oss-120b at $0.35 / $0.75 and claims ~3,000 tokens/s (a vendor claim); Artificial Analysis medians over the 72 hours before 2026-09-22 measured Cerebras at 1,744.5 tokens/s vs Groq's 476.3.
- Price floor: DeepInfra listed gpt-oss-120b at $0.037 / $0.17 per million tokens on 2026-09-22 — roughly a quarter of Groq's input rate on the same checkpoint — with a "zero retention" self-label.
- Third-party spread: across 18 providers serving gpt-oss-120b, blended prices varied up to 8.9× (from $0.04 on CoreWeave to $0.39 on Cloudflare), per Artificial Analysis.
What you're actually replacing Groq for
"Groq alternative" bundles three different migrations, and they point at different providers:
- Model coverage. Groq's public catalog is short by design: a handful of production models (gpt-oss-120b, gpt-oss-20b, Qwen3.8-27B, Whisper) plus enterprise and preview rows. If the checkpoint you need — Kimi K3, DeepSeek V4 Flash, GLM-5.3 — isn't on Groq's list, no amount of LPU speed fixes that.
- Per-token price. Groq's published rates are mid-pack, not bottom-of-market, for the models it does serve.
- Rate limits and enterprise terms. Groq's developer plan caps gpt-oss-120b at 250K tokens-per-minute and 1K requests-per-minute, and the Llama tiers are sales-led. Teams that hit those ceilings shop for dedicated capacity elsewhere.
Decide which of the three is your actual problem before comparing the table below — a cheaper per-token host and a broader catalog solve different pains.
Groq's list price, as a baseline
From Groq's supported-models page, observed 2026-09-22: gpt-oss-120b at $0.15 input / $0.60 output per million tokens (listed speed ~500 tokens/s, 131K context, developer rate limits of 250K TPM / 1K RPM); gpt-oss-20b at $0.075 / $0.30 (~1,000 tokens/s); qwen3.8-27b at $0.80 / $4.00 (~450 tokens/s). Llama 3.1 8B and Llama 3.3 70B are enterprise rows without public prices. Groq also markets LPX working alongside NVIDIA GPUs, positioning "fast or affordable is no longer a tradeoff" — a vendor claim.
The 8 Groq alternatives, compared
Sorted by use case, not by a claimed ranking. Prices are list rates observed 2026-09-22 and move weekly — confirm on each provider's page before budgeting. Where a provider does not publish a per-model list, we say so rather than invent a number.
| Provider | Serving approach | gpt-oss-120b list (in/out, $/M) | Speed signal | Privacy posture |
|---|---|---|---|---|
| Cerebras | Wafer-scale WSE silicon | $0.35 / $0.75 | ~3,000 t/s claimed; 1,744.5 t/s measured (AA) | Centralized; self-published terms |
| SambaNova | Custom RDU silicon | $0.22 / $0.59 | 225 t/s via OpenRouter routing; 98% endpoint accuracy (AA) | Centralized; self-published terms |
| DeepInfra | GPU serverless | $0.037 / $0.17 | 352 t/s (Turbo tier, AA median) | "Zero retention" self-label on this model |
| Novita | GPU serverless | ~$0.05 / $0.25 (AA) | 131 t/s (AA) | Centralized; self-published terms |
| Fireworks AI | GPU serverless + dedicated | Not on Groq's overlap; Kimi K3 $3.00 / $15.00 | — | Centralized; self-published terms |
| Nebius | GPU cloud (Token Factory) | Blended $0.20 (AA) | 295 t/s measured (AA) | EU/US regions; self-published terms |
| Chutes | Decentralized, TEE-attested | Different catalog; Qwen3-32B-TEE $0.104 / $0.416 (2026-09-15) | — | TEE-attested; "operators can't see prompts" is a claim |
| OpenRouter | Aggregator router | $0.15 / $0.60 listed 2026-09-21 | Varies by backend | Retention varies by underlying provider |
1. Cerebras
Cerebras is the direct answer to "I like Groq's speed model but need more of it." It serves a similarly short list on wafer-scale silicon, and its pricing page lists gpt-oss-120b at $0.35 / $0.75 per million tokens at a claimed ~3,000 tokens/s, with Qwen 3.8 27B at $0.99 / $1.49 (~1,850 tokens/s claimed). Independent medians from Artificial Analysis back up the direction: Cerebras topped both output speed (1,744.5 t/s) and time-to-first-token (1.64s) among the 18 providers benchmarked, at a blended $0.39 per million tokens — the priciest blend in the field. You pay a premium for the fastest stream; whether that premium is worth it depends entirely on whether latency is your product requirement.
2. SambaNova
SambaNova is the other custom-silicon house: its RDU (Reconfigurable Dataflow Unit) hardware serves gpt-oss-120b at $0.22 / $0.59 per million tokens — slightly under Groq's $0.15 / $0.60 on input-weighted bills but with a cheaper output rate than Cerebras. Its published board also covers MiniMax-M2.7 ($0.60 / $2.40), Llama 3.3 70B ($0.60 / $1.20), and gemma-4-31B — checkpoints Groq either enterprise-quotes or doesn't carry. On OpenRouter's routing, SambaNova's gpt-oss-120b endpoint ran at 225 tokens/s with 97.5% uptime. SambaNova positions itself at enterprise and sovereign deployments, so treat the self-serve board as the entry tier, not the whole story.
3. DeepInfra
DeepInfra is the budget play, and on the one checkpoint where a direct comparison is possible it is not close: gpt-oss-120b at $0.037 / $0.17 per million tokens as of 2026-09-22 — roughly 4× cheaper on input and 3.5× cheaper on output than Groq's list. Its model page carries a "zero retention" label for this endpoint (a self-published claim, not an independent verification) and offers private dedicated deployments. Artificial Analysis measured DeepInfra's standard endpoint as the #2 cheapest blended ($0.05/M) but slow (~46 t/s); its Turbo tier hits 352 t/s. If your workload is throughput-tolerant — batch processing, offline evals, background summarization — DeepInfra is the first row to price against Groq.
4. Novita
Novita is the other budget serverless row in Artificial Analysis's gpt-oss-120b field: $0.05 blended per million tokens (third cheapest of 18, behind CoreWeave and DeepInfra), at 131 t/s measured with function calling and JSON mode supported. It does not publish a single cross-catalog price list DeAI could rank, so pull the exact model you need from its site. Treat it the way you'd treat any fast-repricing budget host: confirm the rate on the day you buy, and benchmark endpoint quality on your own evals — Artificial Analysis scored its endpoint accuracy at 91% of reference, mid-pack in the field.
5. Fireworks AI
Fireworks AI is the move when the problem is catalog breadth rather than the specific checkpoint. Its serverless docs list per-token rates for models Groq doesn't serve — Kimi K3 at $3.00 / $15.00, DeepSeek V4.1 Flash at $0.30 / $1.20, GLM-5.3 at $1.40 / $4.40 (Standard tier, cached-input rates cheaper, observed 2026-09-22) — plus US-only variants at a 1.5× premium and batch inference at half price. On-demand dedicated deployments run $8/hour for an H100 up to $20/hour for a GB300. It was not among the 18 providers Artificial Analysis tracked for gpt-oss-120b on the observation date, so there is no like-for-like speed row here — Fireworks competes on catalog and tooling (LoRA and full fine-tuning), not on this benchmark. For a head-to-head with its closest rival, see Together AI vs Fireworks AI.
6. Nebius
Nebius is the GPU-cloud route with a per-token API (Token Factory) alongside raw GPU rental, and EU plus US regions that matter for data-residency requirements. In Artificial Analysis's gpt-oss-120b field it measured 295 t/s with 97% endpoint accuracy at a $0.20 blended price — a solid middle of the speed and price distributions, notably faster than the budget hosts. Its public prices page is GPU-hour oriented, so confirm the per-token rate for your model before treating it as a drop-in. Choose Nebius when EU residency or the option to fall back to dedicated GPUs matters as much as the API rate.
7. Chutes (Bittensor SN64)
Chutes is the decentralized, confidentiality-first option: a serverless inference network on Bittensor serving open models in hardware-attested TEEs, where it claims prompts are invisible to GPU operators — an attestation/policy claim DeAI has not independently verified. Its catalog does not overlap Groq's headline models; observed prices from 2026-09-15 include Qwen/Qwen3-32B-TEE at $0.104 / $0.416 per million tokens, moonshotai/Kimi-K3-TEE at $3.00 / $15.00, and deepseek-ai/DeepSeek-V4-Flash-0731-TEE at $0.440 / $1.320. It is not a speed play and not a price play — it is the pick when verifiable execution, not throughput, is the requirement. For the broader category, see the decentralized AI inference networks that actually work.
8. OpenRouter (the aggregator route)
OpenRouter is not a host but a router across many of them, and it is the fastest way to comparison-shop Groq's exact checkpoints. Its catalog listed openai/gpt-oss-120b at $0.15 / $0.60 per million tokens on 2026-09-21 — matching Groq's list that week, after a low of $0.037 / $0.17 on 2026-09-07 that DeAI's Price Index delta flagged as a possible catalog data revision — with 445 models in the catalog overall. The catch: the listed rate is what the router charged for that slug that day, it can include routing markup or a cheaper backend, and retention varies by the underlying provider. For the full field see OpenRouter alternatives and DeAI's dated Price Index.
Speed and price, side by side
For the one checkpoint where third-party data exists across hosts — gpt-oss-120b at reasoning effort "high" — Artificial Analysis published these medians (72 hours before 2026-09-22, 10,000-token prompts):
| Provider | Output speed (t/s) | Time to first token (s) | Blended price ($/M, 7:2:1) | Endpoint accuracy |
|---|---|---|---|---|
| Cerebras | 1,744.5 | 1.64 | $0.39 | 87% |
| Groq | 476.3 | 4.90 | $0.14 | 86% |
| DeepInfra (Turbo) | 352.0 | 6.43 | ~$0.20 | 84% |
| Azure | 296.7 | 7.57 | $0.20 | 93% |
| Nebius | 295.5 | 12.4 | $0.20 | 97% |
| Novita | 131 | 0.95 | $0.07 | 91% |
| DeepInfra (standard) | 46 | 0.65 | $0.05 | 97% |
| SambaNova | — | 21.9 | $0.30 | 98% |
Two readings matter. First, endpoint accuracy is not uniform: the same open weights can lose accuracy to quantization or serving configuration, and SambaNova, Parasail, and Amazon Bedrock scored at or above 98% of reference while several faster endpoints scored lower. Second, the price-speed frontier is steep: Cerebras is roughly 3.7× Groq's measured speed at roughly 2.8× its blended price, while the cheapest hosts trade 10–37× slower streams for 3–9× lower cost. Artificial Analysis is a third party; its numbers are point-in-time medians, not a guarantee for your traffic.
The worked math: 1B input / 250M output
At list rates on gpt-oss-120b, a month running one billion input and 250 million output tokens costs:
- Groq: $150 + $150 = $300
- Cerebras: $350 + $187.50 = $537.50
- SambaNova: $220 + $147.50 = $367.50
- DeepInfra: $37 + $42.50 = $79.50
That is a 3.8× spread across first-party list prices for the identical open-weight checkpoint — before prompt caching (which Fireworks prices at a 90% input discount and Groq applies a 50% cache-hit discount to per Artificial Analysis's cache-discount table) and batch tiers change the picture again. If your traffic is cache-heavy, rerun this table with cache rates before concluding anything.
How to choose (the short version)
- You need a model Groq doesn't serve: check Fireworks, Nebius, or OpenRouter's catalog first — coverage, not speed, is the constraint.
- You need Groq-class speed on the same checkpoints: benchmark Cerebras; it is the only host measured meaningfully faster on gpt-oss-120b.
- Throughput-tolerant batch work: price DeepInfra and Novita; the blended-price leaders.
- Prompts must be confidential by construction: Chutes and the TEE-attested tier — read the attestation model yourself.
- You hit Groq's rate limits: dedicated GPU deployments at Fireworks (from $8/GPU-hour) or Nebius, or OpenRouter as a burst-absorber.
Switching is usually a base-URL swap
Groq and most alternatives here expose OpenAI-compatible chat-completions endpoints, which makes comparison-shopping operationally cheap:
from openai import OpenAI
client = OpenAI(
base_url="https://api.deepinfra.com/v1/openai", # swap per provider
api_key=os.environ["PROVIDER_API_KEY"],
)
resp = client.chat.completions.create(
model="openai/gpt-oss-120b", # exact slug varies by provider
messages=[{"role": "user", "content": "ping"}],
)
Keep the base URL and model name in config, not code, and a provider change becomes a deploy-time decision rather than a migration. Three cautions: model slugs are not portable across providers (Groq's llama-3.3-70b-versatile is SambaNova's Meta-Llama-3.3-70B-Instruct); sampling defaults and quantization differ per endpoint, which shows up in eval scores, not price; and a cheaper model that fails your evals is not cheaper after retries. For the mechanics see how to migrate off the OpenAI API and what an OpenAI-compatible API actually is.
Current state (September 2026)
Fast silicon is no longer a two-horse race: Groq now markets LPX working alongside NVIDIA GPUs, Cerebras holds the measured speed crown on the shared open checkpoints, and SambaNova is courting enterprise and sovereign deployments with its RDU line — while commodity GPU hosts keep undercutting all of them on price for throughput-tolerant work. DeAI's Price Index publishes a weekly snapshot precisely because these lists move weekly; OpenRouter's gpt-oss-120b row swung from $0.037 to $0.15 input within two weeks this month. The durable habit is to re-shop quarterly against a dated benchmark rather than locking to a provider on a price you saw once. For the same menu from the other incumbents' side, see the best Together AI alternatives and Fireworks AI alternatives.
FAQ
What is the best alternative to Groq?
There is no single best one. Cerebras is the closest speed-for-speed rival on the open checkpoints Groq serves. DeepInfra and Novita are the budget serverless options. SambaNova is the other custom-silicon cloud with an enterprise focus. Fireworks and Nebius offer the broadest open-model catalogs. The right pick depends on whether you are replacing Groq for model coverage, price, or rate limits.
Is Cerebras faster than Groq?
On the shared gpt-oss-120b checkpoint, Cerebras claims ~3,000 output tokens per second on its pricing page, and Artificial Analysis's independent medians (72 hours before 2026-09-22) measured Cerebras at 1,744.5 tokens/s versus Groq at 476.3 tokens/s, with Cerebras also fastest to first token (1.64s vs 4.90s). Speeds vary with load and model — benchmark both endpoints on your own traffic before committing.
Which Groq alternative is cheapest?
On 2026-09-22 list prices for gpt-oss-120b: DeepInfra listed $0.037 input / $0.17 output per million tokens, the lowest first-party row DeAI verified. Artificial Analysis's blended rankings (7:2:1 cache-input-output) put CoreWeave at $0.04, DeepInfra at $0.05, and Novita at $0.07 per million tokens blended, versus Groq at $0.14. The blend assumes a cache-heavy mix; price your own input/output split.
Do these Groq alternatives keep my prompts private?
Most are centralized clouds that process prompts under self-published retention policies: Groq, Cerebras, SambaNova, DeepInfra, Fireworks, and Nebius each publish their own terms. DeepInfra labels its gpt-oss-120b endpoint "zero retention" — a self-published label, not an independent verification. If prompts must stay confidential by construction, Chutes serves open models in hardware-attested TEEs, though "operators can't see prompts" remains a claim DeAI has not verified.
Can I switch from Groq without rewriting code?
Usually yes. Groq and most alternatives expose OpenAI-compatible chat-completions endpoints, so switching is typically a base-URL and API-key change. Model slugs are not portable — Groq uses llama-3.3-70b-versatile while SambaNova uses Meta-Llama-3.3-70B-Instruct for the same checkpoint — so map exact names before you cut over.
Why does Groq quote "contact sales" for Llama models?
On Groq's model list as of 2026-09-22, Llama 3.1 8B Instant and Llama 3.3 70B Versatile are enterprise-tier rows with no published per-token price; only gpt-oss-120b, gpt-oss-20b, and Qwen3.8-27B carry public self-serve rates. SambaNova, by contrast, lists Llama 3.3 70B at $0.60 / $1.20 per million tokens publicly. If a public Llama price is what you need, that asymmetry is the shortlist.
Questions
- What is the best alternative to Groq?
- There is no single best one. Cerebras is the closest speed-for-speed rival on the open checkpoints Groq serves. DeepInfra and Novita are the budget serverless options. SambaNova is the other custom-silicon cloud with an enterprise focus. Fireworks and Nebius offer the broadest open-model catalogs. The right pick depends on whether you are replacing Groq for model coverage, price, or rate limits.
- Is Cerebras faster than Groq?
- On the shared gpt-oss-120b checkpoint, Cerebras claims ~3,000 output tokens per second on its pricing page, and Artificial Analysis's independent medians (72 hours before 2026-09-22) measured Cerebras at 1,744.5 tokens/s versus Groq at 476.3 tokens/s, with Cerebras also fastest to first token (1.64s vs 4.90s). Speeds move with load and model — benchmark both endpoints on your own traffic before committing.
- Which Groq alternative is cheapest?
- On 2026-09-22 list prices for gpt-oss-120b: DeepInfra listed $0.037 input / $0.17 output per million tokens — the lowest first-party row DeAI verified. Artificial Analysis's blended rankings (7:2:1 cache-input-output) put CoreWeave at $0.04, DeepInfra at $0.05, and Novita at $0.07 per million tokens blended, versus Groq at $0.14. The blend assumes a cache-heavy mix; price your own input/output split.
- Do these Groq alternatives keep my prompts private?
- Most are centralized clouds that process prompts under self-published retention terms: Groq, Cerebras, SambaNova, DeepInfra, Fireworks, and Nebius all publish their own policies. DeepInfra labels its gpt-oss-120b endpoint 'zero retention' — a self-published label, not an independent verification. If prompts must stay confidential by construction, Chutes serves open models in hardware-attested TEEs, though 'operators can't see prompts' remains an attestation claim DeAI has not verified.
- Can I switch from Groq without rewriting code?
- Usually yes. Groq and most alternatives expose OpenAI-compatible chat-completions endpoints, so switching is typically a base-URL and API-key change. Model slugs are not portable — Groq uses llama-3.3-70b-versatile while SambaNova uses Meta-Llama-3.3-70B-Instruct for the same checkpoint — so map exact names before you cut over.
- Why does Groq quote 'contact sales' for Llama models?
- On Groq's model list as of 2026-09-22, Llama 3.1 8B Instant and Llama 3.3 70B Versatile are enterprise-tier models with no published per-token price — speed is documented but pricing requires contacting sales. Only gpt-oss-120b, gpt-oss-20b, and Qwen3.8-27B carry public self-serve rates. That gap is itself a reason builders comparison-shop the same checkpoints elsewhere.
Sources
- Groq Supported Models (GroqCloud docs) — Groq
- Cerebras Inference Pricing — Cerebras
- Cerebras Launches OpenAI's gpt-oss-120B at 3,000 tokens/sec — Cerebras
- SambaNova Cloud Pricing — SambaNova
- DeepInfra — gpt-oss-120b model page — DeepInfra
- Fireworks AI Serverless Pricing (docs) — Fireworks AI
- Nebius AI Cloud Pricing — Nebius
- Chutes API — chutes list — Chutes
- OpenRouter API — list models — OpenRouter
- Artificial Analysis — gpt-oss-120b (high) provider benchmarking — Artificial Analysis
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
