Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Open-Weights Releases

This Week in DeAI: The Open-Weights Supply Wave Gets a Date

Mistral, Reflection, and Aleph Alpha all promised weights this October; Ecosia already switched, and the deprecation wave shows why owning the file matters.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single tall server rack in a cool, dim European data hall, one amber status light glowing against dark metal panels, depicting the week's lead story: Mistral Large 4, a trillion-parameter model in preview whose open weights are promised by the end of October. Illustration: DeAI
A single tall server rack in a cool, dim European data hall, one amber status light glowing against dark metal panels, depicting the week's lead story: Mistral Large 4, a trillion-parameter model in preview whose open weights are promised by the end of October. Illustration: DeAI

Seven days, one pattern with a deadline attached: between October 3 and October 9, every serious contender in the open-weights race either shipped weights, promised weights with a date, or got dropped by a customer who was tired of waiting for them. Mistral put its trillion-parameter Large 4 into public preview and pinned the open-weights release to the end of October. Reflection AI — Nvidia-backed, $8 billion valued — launched Beam, a 501B-parameter MoE pitched not on beating Kimi K3 but on matching GLM-5.2 while using 3-4x less inference compute, with Apache 2.0 weights due later this month. Aleph Alpha simply shipped: Kolibri-1, 78B parameters under Apache 2.0, downloadable in full since October 3. And Ecosia, the Berlin search engine with a European-first brand, stopped waiting altogether and moved to Chinese open weights on EU-hosted Melious — claiming costs roughly halved and calling Mistral "a year behind."

The theme that organizes all of it is the gap between the announcement and the artifact. Three of the week's four biggest model stories are promissory notes: a preview API that is not weights, a launch post with a waitlist instead of a Hugging Face repo, a rumor with no date. Only Kolibri-1 and a handful of smaller releases convert the promise into something you can download and run today. Meanwhile the week's infrastructure news — four hyperscalers retiring hosted open-model endpoints inside one mid-October window, and OpenRouter list prices moving double digits in both directions — is a live demonstration of why the gap matters. The file you own does not get deprecated; the endpoint you rent does. This is the week the supply side of open weights made its biggest promises of the quarter, and the week the demand side showed it has stopped waiting for them.

Key facts

  • Mistral Large 4 entered public preview October 6: a 1T-parameter MoE with 49B active, scoring 38 on the Artificial Analysis index — up from 9 for Mistral Large 3, but behind Claude Opus 5.5 at 58 and behind China's open-weight DeepSeek 4.1 Flash. Open weights are promised by end of October; until then it is API-only at $0.68/$2.09 per million tokens, with the marketing page listing doubled figures and no statement on which set survives the weights release.
  • Reflection AI launched Beam on October 5: 501B total, 23B active, pretrained on 23.8 trillion tokens, RL-refined on 10,500 Nvidia GB300 GPUs. Its own benchmark table puts Beam behind Kimi K3 and DeepSeek V4.1 Flash on Terminal-Bench v2.1 (80.1 vs 88.3 and 90.6), and the headline 3-4x efficiency claim is an estimated-FLOPs calculation whose methodology note excludes prefill, attention, and serving overhead. No weights exist yet; access today is a waitlist.
  • Aleph Alpha shipped Kolibri-1 on October 3: 78B total, 3.46B active, 1M-token context, Apache 2.0, fully downloadable, serving floor of two A100 80GBs or one H200. Benchmarks — 96.9 AIME 2025, 66.4 SWE-Bench Verified — are self-reported against a comparison field a model generation old.
  • Ecosia CEO Christian Kroll told POLITICO Europe that Mistral was "a year behind," that Ecosia was "simply too large a customer" for Mistral's servers, and that the switch to open weights on EU-hosted Melious has "roughly cut our costs in half while improving quality" — an unverified customer claim, made by a buyer whose users include Germany's environment ministry and the UK's NHS.
  • The October deprecation wave: Google retires all 16 open-model Vertex endpoints October 21 (recommending self-deployment on Model Garden for every one), OpenAI shuts 12 legacy GPT snapshots plus five fine-tune lines October 23, Azure already retired Kimi-K2.7-Code October 3, and Snowflake ends two more models October 14.
  • The Price Index delta recorded DeepSeek V4.1 Flash's OpenRouter row doubling to $0.30/$1.20 (the official peak rate, after Fireworks raised its own DeepSeek pricing October 1), Kimi K3 splitting into a 21x input/output spread, and DeepSeek V4 Pro falling 78% on one row while a sibling row lists $4.50/$5.00.
  • CrowdStrike attributed the South Korean bank intrusions to tooling built on ARTEX, an open-source pentest agent, with DeepSeek V4.1-Flash as primary LLM backend — about 66,000 individuals' records across seven firms, with a human directing every step and the attacker's own Claude Code logs providing the forensics.
  • Nous Research confirmed a $90M Series B at a $1.5B valuation to build a paid enterprise version of the open-source Hermes Agent; its 24-million-clone and ~2.5%-of-global-tokens figures are company self-reports (our coverage).

The promised weights: three labs, three different promissory notes

The week's supply stories are not equivalent, and the differences are the diligence. Mistral's Large 4 preview is the most substantive of the three: a real model serving real traffic at published rates, with a generational benchmark jump (9 → 38 on the AA index) and a date attached to the weights — end of October, with a Hugging Face countdown pointing at October 31. But two qualifications travel with it. The first is the benchmark ceiling: even Mistral's own charts show GPT-6 Astra and Qwen3.8 ahead on SciCode-Verified, and on the AutomationBench agentic benchmark GLM-5.3 scores 62.2 against ML4's 59.9 — the weights, when they land, are joining a field that has kept moving. The second is the pricing ambiguity: the docs list $0.68/$2.09, the marketing page lists $1.36/$4.18, and nothing states which survives the weights release. The cyber-benchmark framing deserves its own qualifier — Mistral's 82% and 93% cybersecurity scores measure provider refusal policy as much as capability, since the closed models it beats mostly refuse the task.

Reflection's Beam is a different kind of promise: not a preview but a launch post about a model nobody can run. The architecture story is coherent — 501B total, 23B active, efficiency pitched against the Chinese cluster rather than capability, which the launch post itself concedes ("frontier open models like Kimi K3 remain ahead on raw capability"). But every number is self-reported, TechCrunch reports third parties have already flagged a scoring discrepancy in the table, and the 3-4x efficiency claim is a FLOPs estimate whose own methodology note excludes prefill and serving overhead — the costs that dominate real deployments. The Axios report three days earlier (our PULSE coverage) had framed the same company's entry as imminent with no date; the launch converted rumor into roadmap, but the artifact is still a waitlist. What would convert the claim: a Hugging Face repo, a technical report, and an independent run — the same three things that separated last week's adoption facts from its claims.

Aleph Alpha's Kolibri-1 is the control case: the only flagship-scale European release of the week where the download exists. The details are unusually complete — 78B total/3.46B active, 1,048,576-token context, FP8 weights at roughly 78 GB, a serving floor of one H200, training on 768 B200s in German and Finnish infrastructure, and a public training-data summary filed under the EU GPAI Code of Practice template. The honest caveats are on the capability side: the benchmarks are self-reported, and the comparison field is a model generation old — nobody at Aleph Alpha claims Kolibri-1 is chasing Kimi K3. Its relevance is different: it is proof that "sovereign European open weights" can mean a finished, licensed, downloadable artifact rather than a sovereignty press release, at the exact moment Ecosia's decision showed a European buyer choosing Chinese weights over the European flagship. Kolibri-1 is what Mistral's October 31 and Reflection's "later this month" both have to become.

The buyer who didn't wait: Ecosia, Melious, and the repriced middle of the market

The week's most strategically loaded story was not a launch. Ecosia — the Berlin search engine whose brand is built on independence from non-European tech, which had already moved from Google to OpenAI to Mistral — told POLITICO Europe it is dropping Mistral for open-weight models hosted by Melious, a German platform serving open weights from servers in eight EU countries. Kroll's stated reasons are three: Mistral's models are "a year behind" the competition; Ecosia was "simply too large a customer" for Mistral's overloaded servers; and Mistral's reliance on international investors made it, in his view, "not truly sovereign." The replacement models are Chinese families — Qwen, GLM, Kimi — running on EU soil. Our full coverage has the details; the structural read belongs in this week's frame.

The inversion is the point. The conventional sovereignty argument runs: open weights are the hedge, the domestic lab is the preference. A European institution-serving buyer with a European-first brand looked at the European flagship and chose Chinese open weights on European hosting instead — on cost, quality, and reliability grounds rather than political ones. The timing sharpens it: the departure landed the same week Mistral released Large 4, and Mistral's chief scientist Guillaume Lample responded with a public pitch — "I would encourage them to test our new model as soon as possible. We can even give them early access today if they want." If Large 4's open weights land on schedule at the end of October, an Ecosia pilot would be the cheapest possible reversal of the story; that is the cleanest test of whether "a year behind" was a snapshot or a trajectory.

Two disciplines apply before generalizing. First, the headline saving — "roughly cut our costs in half while improving quality and performance" — is the customer's own characterization, with no independent evaluation of Ecosia's quality or Melious's pricing published, and Melious itself is the weakly documented link: no independent verification of its hosting footprint or model list. Second, the open-weight route carries a known objection that the coverage of this story keeps flattening: NewsGuard's investigation found the five leading Chinese-backed AI models repeatedly failing to correct pro-China falsehoods, and Germany's AI Association draws the technically precise line — overt censorship is largely fixable through targeted retraining, while subtle bias embedded in training data is much harder to remove. Both positions can be true at once. What Ecosia's decision proves is narrower and more important: for a production workload with institutional users, the weights-plus-competent-hosting stack now beats renting the local flagship's API on the buyer's own scorecard — and that scorecard, not any lab's brand, is what the supply wave is competing for.

What you can actually download this week

Against the three promissory notes, the week's shipped artifacts form their own ledger. Kolibri-1 is the flagship-scale one: full Apache 2.0 weights on Hugging Face since October 3, runnable on a single H200, with the licensing and training-data disclosures that decide whether "open" is reusable or merely downloadable. Z.ai's GLM-5.3 got its own confirmation this week — the full 753B FP8 weights are public and ungated on Hugging Face, 1.42M downloads, and the repo history shows they went up weeks before the "two weeks" framing circulated, under a custom license worth reading before depending on it. The smaller end of the ledger moved too: Liquid AI shipped open weights for d1-3B and an experimental 600M d1-omni decision models that answer in a single forward pass with zero generated tokens, and Google shipped EmbeddingGemma 2, a 740M Apache-2.0 multimodal embedder with Matryoshka truncation to 128 dimensions for up to 6x storage savings.

The pattern across both ledgers is that the category boundary is moving down the stack. Decision models — routing, classification, moderation — are becoming downloadable weights that run on edge hardware, following Fastino's 340M GLiNER2.5-Decide; embedding models are going Apache-2.0 and on-device; and the frontier-coder tier (GLM-5.3) is already self-hostable for anyone with multi-node hardware. Meanwhile the money pages keep repricing the rental alternative: Anthropic cut Claude Haiku 5.5 by up to 90% to $0.10/$0.50 under 100K tokens — exactly matching OpenAI's GPT-6 Luna at that boundary — and our SambaNova pricing page priced the specialist-silicon tier at $0.22/$0.59 on gpt-oss-120b. The practical question for any workload is no longer "open or closed" but "which tier of the open stack, against which repriced closed tier" — and our self-hosting vs API cost analysis prices exactly that comparison. Our explainer on what open-weight models are remains the reference for the licensing mechanics that separate these releases.

The deprecation wave makes the case for the file

If the supply wave is the carrot, the week's infrastructure news is the stick. Four hyperscalers are retiring hosted model endpoints inside the same mid-October window: Google Cloud takes all 16 open-model Vertex endpoints dark on October 21 — DeepSeek-V3.1/V3.2, GLM 5 and 4.7, gpt-oss-20b, Kimi K2 Thinking, Llama 3.3 70B, MiniMax M2, five Qwen3 variants, and more — with Google's recommended alternative for every single one being self-deployment on Model Garden. OpenAI shuts 12 legacy GPT snapshots plus five fine-tune lines on October 23. Azure already retired Moonshot's Kimi-K2.7-Code on October 3 and ends gpt-4.1-nano October 14; Snowflake Cortex ends claude-4-sonnet and openai-gpt-4.1 the same day. The asymmetry is the lesson: the open weights stay downloadable; the hosted convenience does not. Google's own deprecation page is, read carefully, an argument for the weights.

The pricing layer added its own volatility the same week. The Price Index delta — our fourth standardized observation — caught DeepSeek V4.1 Flash's OpenRouter row doubling to $0.30/$1.20 (the official peak rate, following Fireworks' October 1 DeepSeek price increase), Kimi K3 splitting into a 21x input/output spread ($0.67 in, $14.00 out), DeepSeek V4 Pro falling 78% on one OpenRouter row while a sibling row lists $4.50/$5.00, and gpt-oss-120b reverting to its September level. Two observations in a row have now flagged aggregator rows diverging from official pages — the index exists precisely because the rental layer's prices are neither stable nor always what the official page says. Against that, the ownership path has its own recurring costs, which our cost-curve decomposition puts in context: the fixed-capability price decline runs roughly 47% per quarter, so the arbitrage between owning and renting re-prices itself every quarter, in both directions. The deprecation calendar is the tiebreaker: renting means accepting someone else's retirement schedule, and October 21–23 is what that schedule looks like when four providers land in one window. Our cheapest-LLM-API page tracks the resulting field.

The security ledger: what the downloadable stack costs

The week also supplied the sharpest concrete test yet of the open-weights security argument, and it cut in both directions. CrowdStrike's October 7 report tied the South Korean bank intrusions — at least seven financial firms, roughly 66,000 individuals' records — to an attack stack that was open-weight-first and self-hosted: ARTEX, a free open-source pentest agent, calling DeepSeek V4.1-Flash through a reseller, supplemented by GLM-5.3 and Grok 4.6, orchestrated by Claude Code. X spent three days arguing about whether "AI hacked the banks." The verified record is narrower and more useful: Korea's Financial Security Institute states the AI "did not act independently without human involvement — a hacker used the AI as a tool," the tradecraft was automated known technique rather than invented capability, and the single most consequential detail is that the attacker left his own Claude Code session logs in open directories, handing CrowdStrike the forensics.

Two structural points survive the hype. First, the enforcement surface: when inference runs through a provider or reseller, the provider sees prompts and can enforce policy; when it runs self-hosted on rented GPUs, that surface disappears — which is the dual-use risk Anthropic's GLM-5.3 red-team report framed in the abstract, now with a production example. Second, provenance: ProjectDiscovery's under-$50 backdoor demo showed a Qwen2.5-7B fine-tune — poisoned with 125 rows, trained 2.5 hours on one L4 — that passes evals and then exfiltrates .env credentials on a trigger phrase, with the payload remote: the weights carry only a URL, so one commit swaps the behavior without retraining. Guardrails on open weights are defaults, not locks, and a downloaded model's history is part of its attack surface. The counterweight the same week shipped: our SambaNova pricing page and the Harvard-Chutes dataset — 6.1 billion production serving requests from a live decentralized network, verifiable by download — are reminders that the transparency of the open stack is also its auditability. The lesson is not "open weights are unsafe." It is that owning the stack means owning its verification, which is a feature when you do it and an exposure when you skip it.

The money followed the weights

The week's funding story is the open-stack thesis priced in venture terms. Nous Research — the lab behind the open-source Hermes Agent — raised a $90M Series B at a $1.5B valuation, led by Robot Ventures, to build Hermes for Businesses, a paid enterprise version of an open-source agent. The Wall Street Journal reports the company was at roughly $36 million annualized revenue by mid-September and expects to pass $100 million before year-end; the 24-million-clone count and the claim that Hermes drives about 2.5% of global token usage are company self-reports, and the fundraise note itself says so. What the round signals is a business model that did not exist eighteen months ago: monetizing ownership-of-the-stack rather than renting intelligence — an open-weights-native lab selling exactly the property the deprecation wave is teaching enterprises to want. AkashML's launch-provider role for Architect's per-prompt auction router is a second market-design signal in the same direction: decentralized supply being quoted dynamically, though every volume figure so far is provider self-report.

What the week changed for builders

The convergence, stated plainly: the supply side has dates (Mistral October 31, Reflection later this month), the shipped side is broader than ever (Kolibri-1, GLM-5.3 full weights, two decision models, an embedder), the rental side is repricing and deprecating in the same fortnight, and the security ledger now has a concrete production incident on both sides of the openness line. Five moves follow:

  • Treat every promised weights date as a plan, not inventory. Size nothing around October 31 or "later this month" until a Hugging Face repo exists. The checklist that converts a claim: repo, license, technical report, independent run. Kolibri-1 is what a completed release looks like; Beam is what a pending one looks like.
  • Own the deprecation calendar or accept someone else's. Google's 16 Vertex endpoints die October 21 with self-deployment as the recommended migration — the provider's own escape hatch is the weights. Map every hosted open-model endpoint you use against the October wave and decide per endpoint: migrate to Model Garden, move to a specialist host, or take the weights in-house (the cost comparison).
  • Re-price the rental tier before re-deciding the ownership tier. DeepSeek doubled on one router row, fell 78% on another, and Kimi K3 split 21x in the same snapshot — the Price Index method is the discipline, and the closed side moved too (Haiku 5.5's 90% cut). Any ownership decision made against stale prices is wrong in some direction.
  • Verify what you download. The ARTEX incident and the ProjectDiscovery backdoor are the same lesson at different scales: provenance, refusal configuration, and payload provenance are the operator's job once the weights are yours. Benchmark refusals rather than trusting model cards, and treat any fine-tune of unknown origin as hostile until proven otherwise.
  • Watch the buyer signals, not just the launches. Ecosia's switch is one buyer's unverified cost claim, but it is a procurement decision by an institution-serving European company — the class of decision that propagates. If Melious publishes independent hosting verification, or a second EU buyer follows, the "sovereignty means domestic labs" assumption breaks the way AT&T broke the "enterprise means closed" assumption last week (our previous hub).

The week in numbers

  • 38 (from 9) — Mistral Large 4's Artificial Analysis index score versus Mistral Large 3's, behind Claude Opus 5.5 at 58 (Artificial Analysis via Willison).
  • ~2x — the gap between Mistral's documented preview rates ($0.68/$2.09 per million tokens) and its marketing page's figures; which set survives the weights release is unstated.
  • 501B / 23B — Reflection Beam's total and active parameters, with all benchmarks self-reported and no weights downloadable as of October 9 (Reflection).
  • 78B / 3.46B — Kolibri-1's total and active parameters; Apache 2.0, full weights live since October 3, serving floor one H200 (model card).
  • ~50% — Ecosia's claimed cost reduction from switching to open weights on Melious — the customer's own unverified characterization (POLITICO).
  • 16 / 12 / 2 — hosted open-model endpoints retired October 21 by Google, legacy snapshots October 23 by OpenAI, and models October 14 by Snowflake; Azure already retired Kimi-K2.7-Code October 3.
  • +100% / -61% / +65% — the week's OpenRouter moves: DeepSeek V4.1 Flash input+output doubled to $0.30/$1.20, Kimi K3 input fell to $0.67, Kimi K3 output rose to $14.00 (Price Index delta).
  • 66,000 — individuals' records exposed across seven South Korean financial firms, per Korea Herald figures; the attack stack was open-weight-first and self-hosted, with a human directing every step (GovInfoSecurity).
  • Under $50 — ProjectDiscovery's total cost to build a credential-exfiltration backdoor into a 7B model: 125 poisoned rows, 2.5 hours on one L4, payload hosted remotely (ProjectDiscovery).
  • $90M / $1.5B — Nous Research's Series B and reported valuation, funding a paid enterprise version of an open-source agent; traction figures are self-reports (WSJ).

The spokes: what each story established

  • Mistral Large 4 preview (7 Oct) — the dated promise: 1T-param MoE at a real generational jump, weights pinned to end of October, pricing ambiguity flagged, cyber-benchmark framing qualified.
  • Reflection AI's Beam launch (6 Oct) — the promissory launch: efficiency pitched against the Chinese cluster, self-reported benchmarks with a flagged discrepancy, waitlist instead of weights.
  • The Axios Reflection rumor (5 Oct) — the pre-launch claim: unnamed sources, unverified compute deals, no date — the launch post three days later is how such claims convert.
  • Aleph Alpha's Kolibri-1 (4 Oct) — the completed release: full Apache 2.0 weights, complete disclosures, honest caveat that the comparison field is a generation old.
  • Ecosia's switch to Melious (9 Oct) — the demand-side decision: a European brand buyer choosing Chinese open weights on EU hosting over the European flagship, with the bias objection stated fairly.
  • The October deprecation wave (5 Oct) — the stick: four hyperscalers retiring hosted endpoints in one window, with self-deployment as Google's own recommended exit.
  • The Price Index delta (5 Oct) — the rental layer's volatility: DeepSeek doubled, Kimi K3 split 21x, aggregator rows diverging from official pages.
  • CrowdStrike's ARTEX attribution (9 Oct) — the security record: an open-weight-first, self-hosted attack stack, human-directed, forensics handed over by the attacker's own logs.
  • The ProjectDiscovery backdoor (7 Oct) — the provenance lesson: a $50 poisoned fine-tune that passes evals, with a remotely swappable payload.
  • Nous Research's Series B (8 Oct) — the funding signal: an open-source-native lab monetizing ownership-of-the-stack, with self-reported traction labeled as such.
  • The SambaNova pricing page (6 Oct) — the specialist-silicon tier priced: seven open models on RDU at $0.22/$0.59 up to $3/$4.50.
  • The week-ahead calendar (5 Oct) — the frame the week wrote against: COLM 2026, OpenRouter's State of Models, and every rumored release still undated.
  • The week's daily briefs (3–9 Oct) — the day-by-day record: zkAPI on mainnet, Prime Intellect's GB200 serving, GLM-5.3 weights confirmed, textGrain provenance, EmbeddingGemma 2, the Harvard-Chutes serving dataset.
  • Last week's hub (2 Oct) — the adoption record this week's supply stories answer: AT&T at 40%, Vercel's 56% token majority, and the trust-stack migration.
  • The Saturday roundup (3 Oct) — the prior week's seven stories in shorter form, for the summary-level view.

What to watch

Whether Mistral's October 31 lands — the license decides whether Large 4 is genuinely reusable or open in name, and Artificial Analysis re-scores it the day the weights ship. Whether Reflection's Apache 2.0 release comes with the technical report that would let anyone check the 3-4x efficiency claim, and whether the flagged scoring discrepancy survives independent runs. Whether Ecosia's "costs halved" claim gets an independent evaluation, whether Melious publishes verifiable hosting and model-list documentation, and whether a second EU institutional buyer follows — one buyer is an anecdote; two is a procurement pattern. Whether any other Vertex tenant publicly completes the Model Garden migration before October 21, which would establish the deprecation-to-self-host playbook everyone else will copy. And on the security ledger: whether the Korean National Police Agency's task force produces confirmed attribution, and whether any API reseller faces the model-policy question the ARTEX stack raised — what obligation, if any, a reseller has when one customer account drives pentest agents at intrusion scale.

Questions

When are the Mistral Large 4 open weights coming?
Mistral says the weights, license, and architecture details will be released by the end of October 2026, with a Hugging Face countdown pointing to October 31. Until then the 1T-parameter model is API-only at $0.68/$2.09 per million tokens. A promised date is a plan, not a release.
Is Reflection AI's Beam available to download?
No. Beam was announced October 5 as a 501B-total/23B-active sparse MoE, but as of October 9 there is no Hugging Face repository; the only access path is an early-access waitlist. Apache 2.0 weights, a technical report, and FP8/NVFP4 quantizations are promised later in October, and every benchmark in the launch post is self-reported.
What open-weight models shipped this week that I can download today?
Aleph Alpha's Kolibri-1 (78B-total/3.46B-active English-German MoE, Apache 2.0, full weights on Hugging Face since October 3), Liquid AI's d1-3B and d1-omni-600M edge decision models, Google's EmbeddingGemma 2 (740M, Apache 2.0), and Z.ai's full 753B GLM-5.3 FP8 weights, confirmed live and ungated on October 5.
Why did Ecosia switch from Mistral to open-weight models?
CEO Christian Kroll told POLITICO Europe that Mistral's models were 'a year behind' the competition, that Ecosia was 'simply too large a customer' for Mistral's overloaded servers, and that Mistral's reliance on international investors was 'not truly sovereign.' Ecosia now routes through Melious, a German platform hosting Chinese open-weight families (Qwen, GLM, Kimi) on servers in eight EU countries, and claims costs roughly halved — an unverified customer characterization.
Which hosted open-model endpoints are being retired in October 2026?
Google Cloud retires all 16 open-model Vertex endpoints on October 21, OpenAI shuts 12 legacy GPT snapshots and five fine-tune lines on October 23, Azure retires Kimi-K2.7-Code (done October 3) and gpt-4.1-nano (October 14), and Snowflake ends claude-4-sonnet and openai-gpt-4.1 on October 14. The open weights stay downloadable; the hosted convenience does not.

Sources

  1. Introducing Mistral Large 4 — Mistral AI
  2. Introducing Beam: Reflection's 501B open-weight model — Reflection AI
  3. Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
  4. German search engine ditches Mistral, bets on Chinese open-source AI — POLITICO Europe
  5. CrowdStrike: Unknown Threat Actor Uses ARTEX to Target South Korean Finance — CrowdStrike Intelligence
  6. How abliterated models can get you pwned — ProjectDiscovery
  7. A note on our fundraise — Nous Research
  8. Open model deprecations | Google Cloud — Google Cloud
  9. Model deprecations | OpenAI — OpenAI

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →