Cerebras vs Groq is the matchup between the two best-known non-GPU inference stacks, and on the model both serve openly — OpenAI's gpt-oss-120b — third-party medians on 2026-09-29 put Cerebras well ahead on speed and Groq well ahead on list price. Neither carries a broad catalog, so which one you pick depends as much on what your model is as on what your budget is.
Key facts
- Speed (independent): on gpt-oss-120b (high), Artificial Analysis's 72-hour medians measured Cerebras at 1,782.6 tokens/s vs Groq at 476.2 — about 3.7x — with first-answer-token times of 1.59s vs 4.93s on a 10,000-token input (Artificial Analysis, accessed 2026-09-29).
- List price (same checkpoint): Groq lists gpt-oss-120b at $0.15 in / $0.60 out per million tokens; Cerebras lists it at $0.35 / $0.75 (Groq models, Cerebras pricing, both accessed 2026-09-29).
- Blended cost: on Artificial Analysis's 7:2:1 cache-input-output blend, Groq lands at $0.14 per million tokens vs Cerebras at $0.39 — the highest blended row among the 18 providers tracked (Artificial Analysis).
- Catalog breadth: Groq's self-serve production list carries 3 priced text models (gpt-oss-120b, gpt-oss-20b, qwen3.8-27b) plus an audio model; both of its Llama tiers are enterprise-quoted with no public price. Cerebras's shared self-serve catalog lists 2 models (gpt-oss-120b, Qwen 3.8 27B), with more families on enterprise/dedicated (Groq docs, Cerebras docs).
- Rate limits: Groq's developer plan allows 30 RPM / 8K TPM / 1K RPD on gpt-oss-120b; Cerebras's developer tier allows 1,000 RPM with a 1M uncached TPM cap on the same model (Groq rate limits, Cerebras rate limits).
What each one is, defined
Groq runs open-weight models on its own Language Processing Unit (LPU) — custom silicon with a small amount of on-die SRAM and a deterministic, sequential dataflow pipeline. A single LPU cannot hold a large model's weights, so Groq stitches many LPUs into a fabric; the pitch is high, consistent token throughput on a curated model list. The LPU is inference-only hardware.
Cerebras runs models on its Wafer-Scale Engine (WSE-3): one enormous wafer holding 900,000 cores and 44GB of on-die SRAM connected by an on-wafer fabric, so a model's weights never leave the chip. It supports 16-bit precision natively, sells shared self-serve inference on two models and dedicated enterprise capacity for more, and — like Groq — prices per million tokens on a narrow catalog.
Both solve the same problem (GPU memory-bandwidth bottlenecks during decode) with different mechanics: Groq spreads a workload across many small chips; Cerebras puts it on one giant one. Those two mechanisms show up directly in the throughput and latency numbers below.
Cerebras vs Groq: side by side
All figures dated 2026-09-29; prices and speeds move weekly, so re-verify on the linked sources before you buy.
| Criterion | Cerebras | Groq |
|---|---|---|
| Silicon | Wafer-Scale Engine (WSE-3), 900k cores, 44GB on-die SRAM | LPU, ~230MB on-die SRAM per chip, multi-chip fabric |
| gpt-oss-120b list price | $0.35 in / $0.75 out per M | $0.15 in / $0.60 out per M |
| Qwen 3.8 27B list price | $0.99 / $1.49 per M | $0.80 / $4.00 per M |
| gpt-oss-120b speed (listed) | ~3,000 tokens/s (vendor claim) | ~500 tokens/s (vendor claim) |
| gpt-oss-120b speed (independent median) | 1,782.6 tokens/s | 476.2 tokens/s |
| First-answer token (independent median, 10k input) | 1.59s | 4.93s |
| Self-serve catalog | 2 models shared (gpt-oss-120b, Qwen 3.8 27B); more via enterprise | 3 priced text models + Whisper; Llama tiers enterprise-quoted |
| Rate limits (developer tier, gpt-oss-120b) | 1,000 RPM; 1M uncached TPM (3M total TPM) | 30 RPM; 8K TPM; 1K RPD |
| Caching | Cache hits don't count against uncached TPM; cache-hit row $0.10/M | Cache reads $0.075/M; cached tokens don't count toward rate limits |
| OpenAI-compatible | Yes | Yes |
Where the numbers land:
- Speed: Cerebras leads on every third-party measure on the shared checkpoint — output tokens/s, time-to-first-token, and end-to-end response time for a 500-token reply (1.87s vs Groq's 5.98s median). Cerebras's own materials claim up to ~6x; the independent medians currently say ~3.7x on output speed and more on first-token time. The vendor figure and the independent figure disagree in Cerebras's favor in both cases, which is worth noting: even the independent measurement sits well below the ~3,000 tokens/s Cerebras lists, likely reflecting live load.
- Price: Groq's list rates on gpt-oss-120b are lower on both input and output. On Qwen 3.8 27B the picture flips direction depending on your token mix: Groq is cheaper on input ($0.80 vs $0.99) but much more expensive on output ($4.00 vs $1.49) — a workload that generates long answers pays 2.7x more per output token on Groq's list than on Cerebras's.
- Limits: this is the underappreciated row. Groq's self-serve plan caps a single organization at 8,000 tokens per minute on gpt-oss-120b before enterprise negotiation — enough for an interactive app, not for a busy agent fleet. Cerebras's paid tier starts an order of magnitude higher. If your loop burns tokens continuously, Groq's sticker price may not survive its rate-limit page.
Which matters when
Pick Cerebras when the loop is latency-bound. Voice assistants, interactive agents, and code-completion products live or die on time-to-first-token and tail latency. The independent medians put Cerebras's first answer token at 1.59s vs Groq's 4.93s on a 10k-token input, and its 500-token end-to-end at 1.87s vs 5.98s. If the requirement is "the model must finish talking before the user hangs up," that gap is the product.
Pick Groq when cost dominates and your model is on its list. At $0.15/$0.60 per million tokens on gpt-oss-120b, Groq's list is meaningfully below Cerebras's $0.35/$0.75 — roughly half on output. For queued workloads (batch summarization, content pipelines, evaluation runs) where nobody is watching the stream, throughput per dollar beats time-to-first-token, and Groq's prices are the lower sticker.
Watch the rate-limit pages before you commit. Groq's published self-serve caps (30 RPM, 8K TPM, 1K RPD on gpt-oss-120b) are tight enough that real agent workloads will need the enterprise conversation. Cerebras's developer tier is generous by comparison, but its pricing tier page lists only two models at self-serve rates; anything else means the enterprise quote.
Check the catalog last — it's the binding constraint. Both catalogs are curated for their silicon, and neither serves proprietary frontier models (no GPT, Claude, or Gemini) or the largest open checkpoints (no DeepSeek V4, no Kimi K3). If your stack depends on a specific open model, verify it appears on the provider's current list before benchmarking anything. On gpt-oss-20b — the other checkpoint both serve — Groq lists $0.075/$0.30 at ~1,000 tokens/s; Cerebras's self-serve price sheet does not list it, so on that model Groq is the only self-serve option of the two.
For the wider field beyond these two — including the GPU hosts that serve any model at lower speed — see DeAI's roundup of Groq alternatives, which benchmarks the same checkpoints across 8 providers, and the model-level spread in the Price Index.
The switching cost is a base-URL swap
Both endpoints speak the OpenAI chat-completions schema, so moving between them is a two-line change — though their self-serve rate-limit tiers differ enough that you should re-check the limits page after any switch:
from openai import OpenAI
client = OpenAI(
base_url="https://api.provider.example/v1", # Cerebras or Groq — swap this line
api_key=os.environ["PROVIDER_API_KEY"],
)
resp = client.chat.completions.create(
model="qwen-3.8-27b", # Groq slug: qwen/qwen3.8-27b — slugs differ per provider
messages=[{"role": "user", "content": "ping"}],
)
curl https://api.provider.example/v1/chat/completions \
-H "Authorization: Bearer $PROVIDER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen-3.8-27b", "messages": [{"role": "user", "content": "ping"}]}'
Two cautions when porting between them: model slugs are not portable (qwen/qwen3.8-27b on Groq, qwen-3.8-27b on Cerebras), and rate-limit math is not portable either — Groq's caps are per-request-and-token flat rates, while Cerebras's are a dual-bucket uncached/total-TPM model where cache hits let you stretch the same uncached budget further. If your prompt reuse is high, Cerebras's effective token ceiling stretches more than its uncached headline suggests.
Current state (September 2026)
The fast-inference market is moving on two axes at once. Groq markets an LPX next-generation LPU alongside NVIDIA GPU capacity and says it is building hundreds of megawatts; Cerebras continues to press its wafer-scale speed lead and lists a growing enterprise/dedicated path beyond its two-model self-serve tier. SambaNova's RDU platform is the third custom-silicon contender and, per the same Artificial Analysis medians, currently sits between the two on speed (709.4 tokens/s on gpt-oss-120b at a $0.22/$0.59 list price). Prices, catalogs, and rate limits in this comparison are dated 2026-09-29 and will move; DeAI's Price Index re-shoots the provider table weekly.
Questions
- Is Cerebras faster than Groq?
- On gpt-oss-120b in late September 2026, Artificial Analysis's independent 72-hour medians measured Cerebras at 1,782.6 output tokens per second vs Groq at 476.2 — about 3.7x — and a first-token time of 1.59s vs 4.93s on a 10,000-token input. Speeds move with load; benchmark both on your own traffic.
- Is Groq cheaper than Cerebras?
- On list prices as of 2026-09-29 for the same gpt-oss-120b checkpoint, Groq lists $0.15 input / $0.60 output per million tokens and Cerebras lists $0.35 / $0.75 — Groq's rates are lower on both axes. On Artificial Analysis's 7:2:1 blended measure, Groq lands at $0.14 per million tokens vs Cerebras at $0.39.
- What models do Cerebras and Groq actually serve?
- Both catalogs are narrow. Cerebras's shared self-serve tier lists gpt-oss-120b and Qwen 3.8 27B, with more families on its enterprise/dedicated path. Groq's self-serve catalog on 2026-09-29 lists gpt-oss-120b, gpt-oss-20b, gpt-oss-safeguard-20b, and Qwen 3.8 27B, with Llama 3.1 8B, Llama 3.3 70B, and MiniMax M2.7 enterprise-quoted with no public price. Check each provider's current list before committing.
- Which one should I pick for agents and voice workloads?
- If the loop is latency-bound — voice, interactive agents — third-party medians favor Cerebras on both tokens/s and time-to-first-token, but verify on your own prompts. If per-token cost dominates and your model is on Groq's list, Groq's rates are lower. Neither carries frontier proprietary models; for those you need a GPU cloud.
- Can I switch between Cerebras and Groq without rewriting code?
- Both expose OpenAI-compatible chat-completions endpoints, so switching is a base-URL and API-key change. Model slugs differ slightly (Groq: qwen/qwen3.8-27b; Cerebras: qwen-3.8-27b), so map exact identifiers before cutover, and note Groq's developer-plan rate limits (30 RPM, 8K TPM on gpt-oss-120b as of 2026-09-29) differ from Cerebras's dual-bucket uncached/total token caps.
Sources
- Cerebras Inference Pricing (developer tier) — Cerebras
- Cerebras Model Catalog (shared inference) — Cerebras
- Cerebras Rate Limits (free trial / developer tiers) — Cerebras
- Groq Supported Models (GroqCloud docs) — Groq
- Groq Rate Limits (developer plan) — Groq
- gpt-oss-120b (high) API Provider Benchmarking — Artificial Analysis
- OpenRouter gpt-oss-120b provider table — OpenRouter
- SambaNova Cloud Pricing — SambaNova
- Cerebras CS-3 vs. Groq LPU (vendor comparison) — Cerebras
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
