Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Industry & Business

SambaNova Cloud in 2026: RDU Inference Priced Across Its Whole Catalog

SambaNova Cloud runs open models on RDU silicon. We priced all 7 models as of 2026-10-06: $0.22/$0.59 on gpt-oss-120b up to $3/$4.50 on DeepSeek V3.1.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single wafer-scale compute card — a full silicon wafer densely packed with processor cores — seated upright in an open lab fixture under cool, precise light: the custom dataflow architecture behind SambaNova's RDU-based inference cloud. Illustration: DeAI
A single wafer-scale compute card — a full silicon wafer densely packed with processor cores — seated upright in an open lab fixture under cool, precise light: the custom dataflow architecture behind SambaNova's RDU-based inference cloud. Illustration: DeAI

SambaNova Cloud prices seven open-weight models on its own Reconfigurable Dataflow Unit (RDU) silicon at $0.22 to $3.00 per million input tokens as of 2026-10-06, and its gpt-oss-120b endpoint is the second-fastest third-party-measured route to that checkpoint anywhere. The catch is what surrounds the speed: a seven-model catalog, a 20-million-token daily cap, and no published price on the dedicated tier.

Key facts

  • Price floor: gpt-oss-120b lists at $0.22 input / $0.59 output per million tokens — the lowest row on SambaNova's rate card and the cheapest way to buy that model at its measured speed tier (SambaNova pricing, accessed 2026-10-06).
  • Price ceiling: DeepSeek V3.1 and V3.2 both list at $3.00 / $4.50 — roughly 14x the floor, on a rate card of only seven models (SambaNova pricing).
  • Speed (independent): Artificial Analysis's 72-hour medians on gpt-oss-120b (high) measured SambaNova at 702.5 tokens/s, second only to Cerebras (1,762.9) and ahead of Groq (472.3); time to first answer token is 3.93s vs Cerebras's 1.60s (Artificial Analysis, accessed 2026-10-06).
  • Rate limits: developer-tier accounts are capped at 20 million tokens per day across all models, with 60 RPM / 12,000 RPD on gpt-oss-120b (SambaNova rate limits).
  • Cache discount: MiniMax-M2.7 — the only model with a published cached-input rate — bills cached input at $0.06 per million vs $0.60 uncached, a 90% cut (SambaNova pricing, DeepInfra review).
  • Company footing: SambaNova closed the first tranche of a $1 billion raise at an $11 billion valuation on 2026-07-08 (SambaNova press).

What SambaNova Cloud is

SambaNova Cloud (SambaCloud) is the public, pay-as-you-go inference API of SambaNova Systems, the San Jose startup that has spent a decade building Reconfigurable Dataflow Units — inference-only chips that map a model's computation graph directly onto the processor instead of running it as generic tensor math on a GPU. The current generation, the SN40L, uses a three-tier memory system (on-chip SRAM, 64 GB of HBM, plus DDR DRAM) that lets a single node hold very large models entirely in memory, which is why SambaNova has historically claimed world-record single-stream speeds on big checkpoints like Llama 3.1 405B.

The cloud API speaks the OpenAI chat-completions schema at https://api.sambanova.ai/v1, plus the Anthropic Messages API as of a July 2026 release. The company itself is on firm financial footing: a $1 billion first close at an $11 billion valuation in July 2026, following a $350 million Series E in February, alongside the announcement of the fifth-generation SN50 chip (vendor claims: 5x Blackwell B200 top speed, models up to 10 trillion parameters, 10-million-token context) — but the SN50 ships in the second half of 2026, so today's SambaCloud runs on the prior SN40L generation.

The full rate card, priced (2026-10-06)

Every listed model is open-weight. Prices are SambaNova's first-party list rates, accessed 2026-10-06:

ModelTierContextInput $/MOutput $/MCached input $/M
gpt-oss-120bProduction128k$0.22$0.59—
gemma-4-31B-itPreview128k (multimodal)$0.38$1.15—
Meta-Llama-3.3-70B-InstructProduction128k$0.60$1.20—
MiniMax-M2.7Production192k$0.60$2.40$0.06
MiniMax-M3Preview1M$0.60$2.40—
DeepSeek-V3.1Production128k$3.00$4.50—
DeepSeek-V3.2Preview32k$3.00$4.50—

Three pricing notes the card itself doesn't tell you:

  • The July cut-and-revert. UsagePricing's price history shows SambaNova cut gemma-4-31B-it from $0.38/$1.15 to $0.22/$0.59 in July 2026, then reversed the cut on 2026-08-11, returning the model to its original rate. If you budgeted against the July card, re-check it.
  • The card is narrow and has been narrowing. The rate card carries six to seven models at a time; DeepSeek-V3.1-cb and DeepSeek-R1-Distill-Llama-70B were dropped from the public card in July 2026, and the deprecation log shows Preview models retired with as little as two to three weeks' notice. Anything you pin should be checked against the deprecation record monthly.
  • Cached-input pricing is the exception, not the rule. Only MiniMax-M2.7 carries a cached rate ($0.06/M, 90% below uncached). SambaNova's own measurements claim TTFT drops of 33% on short contexts and 91% at 192k context when caching hits — but if your model isn't MiniMax, you're paying uncached rates on every repeated system prompt.

Speed: second place on the benchmark, first on the rate card's terms

Artificial Analysis benchmarks 18 providers serving gpt-oss-120b (high). SambaNova's 72-hour medians as of 2026-10-06: 702.5 output tokens/s, 3.93s to first answer token, and a 4.64s end-to-end time for a 500-token response — the fastest end-to-end response of the field, because its decode speed more than compensates for a first token that arrives later than Cerebras's. Its Endpoint Accuracy Index scores 98%, tied for the top of the tracked field, which matters: fast endpoints that quantize aggressively can lose accuracy, and this one doesn't appear to.

On Artificial Analysis's 7:2:1 blended price measure, SambaNova's gpt-oss-120b lands at $0.26 per million tokens — more than the batch-tier GPU hosts (CoreWeave $0.04, DeepInfra $0.05) but well under Cerebras ($0.39) for a third of Cerebras's measured speed. The blended math also flatters the GPU hosts, which assume a cache-heavy mix SambaNova's endpoint doesn't currently bill at a cache discount; on a cache-light 1:1 input/output mix, the gap narrows.

On Llama 3.3 70B, DeepInfra's September 2026 industry review measured SambaNova's shared endpoint at 302 output tokens/s with 1.76s TTFT — the highest throughput in its seven-provider table, but against a price of $0.60/$1.20 where DeepInfra's Turbo tier lists $0.10/$0.32 for the same weights at 18 tokens/s. SambaNova's pitch is that the speed is the product; the honest read is that you pay roughly 4–5x the batch-tier floor for it on this model.

Rate limits: the quiet constraint

SambaNova's rate-limit page is the row most comparison charts skip, and it's the one that decides production fit:

  • Developer tier: 20 million tokens per day, account-wide, across all models. A sustained agent loop burning 100k tokens/hour hits the daily wall in under nine days of round-the-clock running — and much faster with parallel traffic.
  • Per-model production caps (developer tier): 60 RPM / 12,000 RPD on gpt-oss-120b, DeepSeek-V3.1, and MiniMax-M2.7; 240 RPM / 48,000 RPD on Llama 3.3 70B.
  • Free tier: 20 RPM, 20 RPD, and 200,000 tokens per day — enough for an afternoon of evaluation, not a pilot. And since August 2026, even the "free" plan requires adding a payment method before you can buy pay-as-you-go credits; the old no-card $5 grant is gone.
  • Above that: contact sales. There is no published price for Enterprise or dedicated capacity, which makes the serverless-to-dedicated crossover impossible to model from the outside.

Contrast: Groq's self-serve developer plan on gpt-oss-120b is 1,000 RPM / 250k TPM; Cerebras's developer tier runs 1,000 RPM with a 1M uncached TPM cap. SambaNova's per-request and daily caps are an order of magnitude tighter. For bursty interactive traffic this rarely binds; for an agent fleet, it is usually the first wall.

Catalog: seven models, no frontier

The full list is four Production models and three Preview models, all open-weight, all dated against the current generation of releases: there is no DeepSeek V4.1 Flash, no Kimi K3, no GLM 5.3 — the checkpoints that dominate October 2026's coverage are absent. Preview models carry short-notice removal risk by policy, and the deprecation record shows SambaNova retires aggressively (Qwen3-32B, DeepSeek-V3.1-Terminus, and the Llama 3.1 8B tier were all deprecated in March–April 2026 with named replacements).

Two things the catalog does uniquely well: MiniMax-M3's 1M-token context (the longest window on this rate card) and gemma-4-31B-it's image-plus-video input, the only multimodal row. If neither matters to you, the catalog question is really "does it have your model at all" — and for most builders building against the October 2026 open-weight frontier, the answer is no.

Privacy: a strong claim, thin documentation

SambaNova's marketing states that SambaCloud "never sees or collects any of your data or user prompts, ensuring full data privacy" — a strong sentence that appears on the product page and in blog posts. DeAI's rule for provider privacy claims is that absence-claims are policy statements, not verified facts, and here the documentation doesn't yet meet the claim: developers in SambaNova's own community forum have asked for explicit, contract-level documentation of how serverless prompts are stored and used, and the public privacy policy we reviewed speaks to corporate personal information rather than making a per-request zero-retention commitment. Treat the "never sees" line as a self-published claim pending written terms — the same standard we apply to every provider in the Provider Trust Tracker. If prompts are sensitive, Chutes-style TEE attestation is a different, verifiable mechanism, and SambaNova does not currently offer it on the public cloud.

SambaNova vs the alternatives, briefly

  • vs Cerebras: the closest architectural rival. Cerebras is roughly 2.5x faster on the shared gpt-oss-120b checkpoint (1,762.9 vs 702.5 tokens/s measured) with a faster first token, and its self-serve catalog is also two models deep; SambaNova's list prices are lower on both axes of that checkpoint ($0.22/$0.59 vs $0.35/$0.75) and its catalog carries larger-context and multimodal rows Cerebras doesn't. For the full head-to-head of the non-GPU incumbents, see our Cerebras vs Groq comparison, which prices the same checkpoints across both.
  • vs Groq: Groq is meaningfully slower on gpt-oss-120b (472.3 tokens/s) but reaches its first token in about the same time (4.95s vs 3.93s — SambaNova still leads, modestly), prices gpt-oss-120b lower ($0.15/$0.60), and grants dramatically higher self-serve rate limits. Groq's Llama tiers are enterprise-quoted, so on Llama 3.3 70B SambaNova is the only one of the two with a public self-serve rate.
  • vs GPU serverless (DeepInfra, Novita, Fireworks, Nebius): those hosts are 3–10x cheaper per token on the same open checkpoints and serve far broader catalogs, at 10–40x lower measured output speed on shared endpoints. If nobody is watching the stream, the GPU hosts win the bill; if your product is the stream, the custom-silicon premium is the point.
  • vs aggregators (OpenRouter): OpenRouter routes to an upstream SambaNova endpoint, so you can A/B the speed through a gateway you already pay for before committing to a first-party account.

Should you use it?

SambaNova Cloud fits when three conditions hold at once: your model is on its seven-row list; your workload is decode-heavy enough that 700 tokens/s beats the GPU hosts' batch-tier economics; and your volume fits under 20 million tokens per day or you're prepared to negotiate Enterprise. Long-document processing on MiniMax-M3's 1M context, voice-adjacent flows on gpt-oss-120b, and multimodal Gemma work are the obvious fits.

It fits poorly when you need a current-frontier checkpoint (DeepSeek V4, Kimi K3, GLM 5.3), when your agents would blow through the daily token cap, or when confidential-compute attestation is a hard requirement — for that, the TEE providers are the category to shop.

For the wider field, DeAI maintains dated roundups of Groq alternatives and Fireworks AI alternatives, and the live price spread across every provider sits in the Price Index. Prices and speeds here are dated 2026-10-06 and will move; this page gets refreshed alongside the weekly Price Index re-shoot.

Switching cost: a base-URL swap

SambaCloud's endpoint is OpenAI-compatible, so moving to it from any OpenAI-shaped client is a two-line change:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.sambanova.ai/v1",
    api_key=os.environ["SAMBANOVA_API_KEY"],
)

resp = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "ping"}],
)
curl https://api.sambanova.ai/v1/chat/completions \
  -H "Authorization: Bearer $SAMBANOVA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "ping"}]}'

Model slugs are not portable across providers (gpt-oss-120b here, openai/gpt-oss-120b on Groq), and reasoning effort is controlled per request on the gpt-oss checkpoints (reasoning_effort: "high"), so map identifiers and sampling defaults before any cutover.

Questions

How much does SambaNova Cloud cost per million tokens?
As of 2026-10-06, SambaCloud's rate card spans $0.22 input / $0.59 output per million tokens on gpt-oss-120b up to $3.00 / $4.50 on DeepSeek V3.1 and V3.2, with Meta-Llama-3.3-70B at $0.60/$1.20, MiniMax-M2.7 at $0.60/$2.40, and Gemma 4 31B at $0.38/$1.15. MiniMax-M2.7 is the only model with a published cached-input rate, $0.06 per million.
Is SambaNova faster than Groq or Cerebras?
On gpt-oss-120b (high) as of 2026-10-06, Artificial Analysis's 72-hour medians measured SambaNova at 702.5 output tokens per second — second only to Cerebras at 1,762.9 and well ahead of Groq at 472.3. On time to first token the order holds (SambaNova 3.93s vs Cerebras 1.60s and Groq 4.95s). Speeds move with load; benchmark on your own traffic.
What models does SambaNova Cloud serve?
Seven, per SambaNova's docs on 2026-10-06: four Production models (MiniMax-M2.7, DeepSeek-V3.1, Meta-Llama-3.3-70B-Instruct, gpt-oss-120b) and three Preview models (MiniMax-M3 at 1M-token context, DeepSeek-V3.2, and gemma-4-31B-it with image and video input). There are no proprietary frontier models, and the newer DeepSeek V4 and Kimi K3 checkpoints are not on the list.
What are SambaNova Cloud's rate limits?
Developer-tier accounts are capped at 20 million tokens per day across all models. Per-model production limits as of 2026-10-06 run 60 RPM / 12,000 RPD on gpt-oss-120b, DeepSeek-V3.1, and MiniMax-M2.7, and 240 RPM / 48,000 RPD on Llama 3.3 70B; free-tier accounts get 20 RPM, 20 RPD, and 200,000 tokens per day. Higher limits require contacting sales.
Does SambaNova train on your prompts?
SambaNova's marketing says SambaCloud 'never sees or collects any of your data or user prompts,' but that is a self-published claim — DeAI has not found a public contractual zero-retention commitment or independent verification behind it. Read SambaNova's privacy policy and ask for written terms if prompts are sensitive.
When does the SN50 RDU ship, and should you wait?
SambaNova announced the fifth-generation SN50 RDU in February 2026 with vendor claims of 5x Blackwell B200 top speed and support for models up to 10 trillion parameters, but DeepInfra's industry review notes SN50 ships in the second half of 2026 — SambaCloud runs the prior SN40L generation today. Buy against measured current-endpoint numbers, not the next chip.

Sources

  1. SambaNova Cloud Pricing (per-token rate card) — SambaNova
  2. SambaNova Cloud Plans (Free / Developer / Enterprise tiers) — SambaNova
  3. SambaCloud Supported Models — SambaNova
  4. SambaNova Rate Limits Policy — SambaNova
  5. SambaCloud Model Deprecations and Migration Guide — SambaNova
  6. gpt-oss-120b (High) API Provider Benchmarking — Artificial Analysis
  7. SambaNova: Models, Performance & Price (Artificial Analysis provider page) — Artificial Analysis
  8. SambaNova Privacy Policy — SambaNova
  9. SambaNova Privacy & data use in developer tier (community thread) — SambaNova Developer Community
  10. Introducing the SN50 RDU: Purpose-Built for Agentic Inference — SambaNova
  11. Best AI Inference Platforms for Speed & Cost in 2026 (DeepInfra industry review) — DeepInfra
  12. SambaNova Cloud Pricing history (UsagePricing blueprint) — UsagePricing
  13. Cerebras Inference Pricing — Cerebras
  14. Groq Supported Models (GroqCloud docs) — Groq
  15. DeepInfra Pricing — DeepInfra

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →