Reflection AI, the Nvidia-backed startup founded by former Google DeepMind researchers, launched Beam on October 5: a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active, pitched on inference efficiency rather than raw capability. Every benchmark is vendor self-reported, and no downloadable weights exist yet.
Key facts
- Reflection's launch post states the architecture: 501 billion total parameters, 23 billion active per token, pretrained on 23.8 trillion tokens.
- The self-reported benchmark table puts Beam at 80.1 on Terminal-Bench v2.1 versus 88.3 for Moonshot AI's Kimi K3 and 90.6 for DeepSeek V4.1 Flash, and 44.4 on DeepSWE v1.1 versus 68.0 (Kimi K3) and 74.2 (DeepSeek).
- The headline efficiency claim is that Beam matches Z.ai's GLM-5.2 on advanced reasoning while using 3-4x less inference compute, and the post's own methodology note says this is an estimated-FLOPs calculation, not measured serving cost.
- The reinforcement-learning run used 10,500 Nvidia GB300 GPUs for four weeks, generating over 100 million rollouts with a maximum context of 256K tokens and roughly 1.3 billion sandboxes.
- Reuters confirms the launch and the framing against Chinese open-weight models such as DeepSeek and Kimi; TechCrunch reports that third parties have already flagged a scoring discrepancy in Reflection's table.
- Weights under Apache 2.0, the technical report, the model card, and FP8/NVFP4 quantizations are promised "later this month." Today the only access path is an early-access waitlist.
What happened
Reflection AI published its launch post on October 5, 2026, and Reuters moved a wire the same day. The company was founded in March 2024 by Misha Laskin and Ioannis Antonoglou, both formerly of Google DeepMind, and raised $2 billion at an $8 billion valuation in October 2025. Beam is text-only, built for coding, reasoning, and agentic workloads, and is undergoing final red-teaming before release.
The numbers come from Reflection, not from anyone else. The launch post's benchmark table puts Beam ahead of Z.ai's GLM-5.2 on several coding rows, including SWE-Bench Pro v2-Hard at 77.2 versus 84.3 for GLM-5.3 and 88.2 for Kimi K3, and roughly level with GLM-5.2 on Terminal-Bench v2.1 at 80.1 versus 81.0. The framing in the post itself is unusual for a launch: Reflection states outright that "frontier open models like Kimi K3 remain ahead on raw capability" and sells Beam's advantage as "efficiency at inference time."
That efficiency claim deserves its own sentence of scrutiny. Reflection says Beam achieves GLM-5.2-comparable reasoning while using 3-4x less inference compute, but the methodology note in the same post says the estimate uses generation forward-pass compute approximated as FLOPs = 2 x active parameter count x mean generated tokens per attempt, using third-party eval data, and explicitly excludes prompt prefill, context-dependent attention operations, and serving overhead. By the company's own description, that is an approximate compute comparison rather than measured inference cost. A real deployment-cost claim would need measured serving data on the hardware builders actually run.
There is also a discrepancy on record already. TechCrunch reports that third parties flagged a scoring difference: Reflection's table lists Qwen's DeepSWE v1.1 result at 51.0, while Qwen's own model card lists 56.6. Reuters, the strongest non-vendor confirmation, verified that the launch happened and reported the headline parameters, but a wire cannot verify a benchmark.
What is unambiguously real is the training scale and what is still missing. The RL campaign ran on 10,500 Nvidia GB300 GPUs over four weeks, sustained an average of 110,000 concurrent rollouts, and drew on a pool of nearly one million mostly synthetic task environments. None of that changes the operative fact for builders: on October 6 there is no Beam checkpoint on Hugging Face, no model card, and no technical report. The only way to use Beam today is to join a waitlist.
Why it matters
The day-one checkpoint supply for frontier-capable open-weight models has run through a short list of labs: Moonshot AI's Kimi, Z.ai's GLM, DeepSeek, and Alibaba's Qwen. Artificial Analysis's independent scoring illustrates the gap Reflection is entering against: Nvidia's Nemotron 3 Ultra, announced in May 2026, was the most intelligent US open-weights model at the time on Artificial Analysis's index at 48, still behind the Chinese-led frontier at Kimi K2.6's 54. A US entrant that ships real weights under a permissive license widens the supply builders can fall back on, which is the whole point of the open-weight release-watch.
The pitch is also a strategic tell. Reflection is not claiming to win the benchmark race; it is claiming a better tokens-per-capability position, aimed at "enterprise coding and agentic workloads." If the claim survives independent testing, it matters for anyone serving open models at volume, because active-parameter count and generated-token count are the two biggest levers in serving economics. If it does not survive, Beam lands as another large model that trails Kimi K3. The claim is checkable only after the weights ship, which is why the October release date is the real news peg.
Builders should also read the packaging. Alongside the model launch, Reuters notes Reflection's earlier-year SpaceX compute deal and its sovereign-AI-factory partnership with South Korea's Shinsegae Group: the same weights are being positioned as something a government or enterprise can run in its own facility rather than only as a hosted endpoint.
Background
This launch closes the loop on a story that started as reporting, not announcement. On October 4, Axios reported, citing unnamed sources, that Reflection was close to releasing a frontier open-weight model pitched as competitive with top Chinese open-weight systems, with large Nvidia-server rental deals at Nebius and SpaceX behind it. DeAI covered that as a PULSE with the explicit note that nothing was verifiable until a Hugging Face repository or announcement post existed. The announcement post now exists; the repository still does not.
The release-watch context matters. DeAI's open-weight model explainer covers why downloadable weights change a builder's risk profile: weights survive any single vendor's pricing or deprecation decision, and licenses decide what you can do with them. Apache 2.0, if it ships as promised, is the most permissive license in common use for models of this scale. Compare that with recent releases DeAI has tracked: Aleph Alpha's Kolibri-1 shipped with weights available and a stated hardware footprint, which is what a verified release looks like, and StepFun's Step 5 Preview was slated to land with both an API and open weights on October 15.
The verification checklist for Beam is now short and concrete: an Apache 2.0 checkpoint on Hugging Face, the technical report, independent scoring from Artificial Analysis or a comparable index, and a re-check of the benchmark table against each lab's own model cards, starting with the DeepSWE discrepancy TechCrunch flagged.
What's next
Three dated or checkable items. First, the weights: Reflection has committed to releasing Apache 2.0 weights, the technical report, the model card, and FP8/NVFP4 quantizations later in October 2026, so the Hugging Face org page is the artifact to watch. Second, independent benchmarks: once weights are out, Artificial Analysis scoring will either support or break the 3-4x efficiency claim, and Qwen's DeepSWE number is the first table entry to re-verify. Third, the access model: today's early-access waitlist is the only way to run Beam, and how quickly Reflection converts waitlist access into downloadable weights will say more about the release than any benchmark row.
Questions
- What is Reflection AI's Beam model?
- Beam is Reflection AI's first open-weight model, announced October 5, 2026: a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, aimed at coding, reasoning, and agentic workloads. It is in final red-teaming, with weights, a technical report, a model card, and FP8/NVFP4 quantizations promised under Apache 2.0 later in October.
- Is Beam better than Kimi K3 or DeepSeek V4.1 Flash?
- On Reflection AI's own benchmark table, no: Beam scores 80.1 on Terminal-Bench v2.1 against Kimi K3's 88.3 and DeepSeek V4.1 Flash's 90.6, and 44.4 on DeepSWE v1.1 against K3's 68.0 and DeepSeek's 74.2. The company's pitch is efficiency, not raw capability. All figures are self-reported, and no independent scoring of Beam exists yet.
- Is the 3-4x inference-efficiency claim verified?
- No. It is a Reflection AI estimate computed as FLOPs = 2 x active parameters x mean generated tokens, using third-party eval data. The company's own methodology note says the estimate excludes prompt prefill, attention, and serving overhead, so it is not a measured deployment-cost comparison. What would verify it: independent runs once weights ship.
- When can I download Beam's weights?
- Reflection AI says the Apache 2.0 weights, technical report, model card, and FP8/NVFP4 quantizations will be released later in October 2026. As of October 6 there is no Hugging Face repository; the only access path today is an early-access waitlist.
Sources
- Introducing Beam: Reflection's 501B open-weight model — Reflection AI
- Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost — TechCrunch
- Nvidia-backed Reflection unveils first AI model to take on Chinese open models (Reuters wire) — Reuters via The News Tribune
- Nemotron 3 Ultra launch announced: high-speed, leading US open weights intelligence — Artificial Analysis
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
