Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Open-Weights Releases

The 8 DeAI stories that mattered this week

Mistral put a date on open weights, Ecosia stopped waiting, four hyperscalers began retiring hosted endpoints, and a $50 backdoor showed what downloads cost.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single tall rack server in a cool, dim European data hall, one amber status light glowing against dark metal panels, depicting the week's lead story: Mistral putting a date on open weights for its trillion-parameter Large 4. Illustration: DeAI
A single tall rack server in a cool, dim European data hall, one amber status light glowing against dark metal panels, depicting the week's lead story: Mistral putting a date on open weights for its trillion-parameter Large 4. Illustration: DeAI

The week open weights stopped being an idea and became a schedule. Mistral put its trillion-parameter Large 4 into public preview and pinned the open-weights release to end of October. Reflection AI — Nvidia-backed, $8 billion valued — launched Beam with a waitlist instead of a Hugging Face repo, while Aleph Alpha quietly shipped Kolibri-1 in full. Ecosia, the Berlin search engine, stopped waiting for any of it and moved to Chinese open weights on EU-hosted Melious. And the infrastructure underneath showed why the file matters more than the endpoint: four hyperscalers began retiring hosted open-model endpoints inside one mid-October window. Eight stories mattered; here is what each one changes.

Key facts

  • End of October — Mistral's promised release window for Large 4's open weights; until then the 1T-parameter MoE is API-only at $0.68/$2.09 per million tokens, scoring 38 on the Artificial Analysis index, up from 9 for Mistral Large 3 but behind Claude Opus 5.5's 58 and behind China's open-weight DeepSeek 4.1 Flash.
  • "A year behind" — Ecosia CEO Christian Kroll's public verdict on Mistral as he moved the company to open weights on EU-hosted Melious, claiming costs roughly halved — an unverified customer characterization.
  • 501B total / 23B active — Reflection AI's Beam, pitched on using 3-4x less inference compute than the Chinese cluster; every benchmark is self-reported and no weights are downloadable yet.
  • 16 / 12 / 2 — hosted open-model endpoints retired October 21 by Google Cloud, legacy snapshots October 23 by OpenAI, and models October 14 by Snowflake; Azure already retired Kimi-K2.7-Code October 3.
  • 66,000 — individuals' records exposed across seven South Korean financial firms, per CrowdStrike's attribution to an attack stack running an open-source pentest agent on DeepSeek V4.1-Flash — with a human directing every step.
  • Under $50 — ProjectDiscovery's total cost to build a credential-exfiltration backdoor into a 7B model that passes evals, then exfiltrates .env files on a trigger phrase.

1. Mistral Large 4 preview: the trillion-parameter promise with a date on it

Mistral's biggest launch yet is also its most conditional. Mistral Large 4 entered public preview October 6 as a 1T-parameter MoE with 49 billion active parameters, scoring 38 on the Artificial Analysis index — a genuine generational jump from Mistral Large 3's 9, but behind Claude Opus 5.5 at 58 and behind China's open-weight DeepSeek 4.1 Flash. Open weights, license, and architecture details are promised by end of October, with a Hugging Face countdown pointing at October 31; until then it is API-only at $0.68/$2.09 per million tokens. The pricing deserves a flag of its own: the docs list those figures, the marketing page lists doubled ones, and nothing states which set survives the weights release. Why it matters: a promised date is a plan, not a release — anyone sizing a self-hosted deployment around a trillion-parameter European candidate should treat October 31 as intent, not inventory. Full story.

2. Ecosia drops Mistral for Chinese open weights on EU-hosted Melious

The demand side stopped waiting. Ecosia — the Berlin search engine whose brand is built on European digital independence — told POLITICO Europe it is replacing Mistral with open-weight models including Qwen, GLM, and Kimi on Melious, a German platform serving open weights from servers in eight EU countries. CEO Christian Kroll called Mistral "a year behind," said Ecosia was "simply too large a customer" for Mistral's overloaded servers, and argued Mistral's reliance on international investors made it "not truly sovereign." The claim that costs "roughly cut in half while improving quality" is the customer's own characterization, not an independent evaluation — and Mistral's chief scientist has publicly invited Ecosia back for early access to Large 4. Why it matters: a European, institution-serving buyer chose Chinese open weights over the European flagship on cost and reliability grounds, not political ones — the clearest signal yet that the weights-plus-competent-hosting stack is beating the domestic-flagship API on the buyer's own scorecard. Full story.

3. Reflection AI launches Beam: 501B parameters, no weights

The most-negotiated launch of the week is the least downloadable one. Reflection AI — Nvidia-backed, valued at $8 billion — introduced Beam on October 5, a sparse mixture-of-experts model with 501 billion total and 23 billion active parameters, pitched on using 3-4x less inference compute than the Chinese open-weight cluster rather than beating it on raw capability. The company's own benchmark table puts Beam behind Kimi K3 and DeepSeek V4.1 Flash on Terminal-Bench v2.1 (80.1 vs 88.3 and 90.6), every number is self-reported, and TechCrunch reports third parties have already flagged a scoring discrepancy in the table. Apache 2.0 weights, a technical report, and FP8/NVFP4 quantizations are promised later in October; access today is an early-access waitlist. Why it matters: the 3-4x efficiency claim is an estimated-FLOPs calculation whose own methodology note excludes prefill, attention, and serving overhead — the costs that dominate real deployments — so treat it as a hypothesis awaiting a downloadable artifact, not a measured result. Full story.

4. Aleph Alpha ships Kolibri-1: the control case for what "shipped" means

The week's only flagship-scale European release where the download actually exists. Aleph Alpha's Kolibri-1 is 78 billion total parameters, 3.46 billion active, with a 1M-token context, under Apache 2.0, fully downloadable since October 3 — full FP8 weights at roughly 78 GB, a serving floor of one H200, training on 768 B200s in German and Finnish infrastructure, and a public training-data summary filed under the EU GPAI Code of Practice template. The honest caveat is on the capability side: the benchmarks — 96.9 AIME 2025, 66.4 SWE-Bench Verified — are self-reported against a comparison field a model generation old, and nobody at Aleph Alpha claims it is chasing Kimi K3. Why it matters: Kolibri-1 is proof that "sovereign European open weights" can mean a finished, licensed, downloadable artifact rather than a sovereignty press release — it is what Mistral's October 31 and Reflection's "later this month" both have to become. Full story.

5. The October deprecation wave: four hyperscalers retire hosted open-model endpoints

The week's infrastructure story is a live demonstration of why items 1 through 4 exist. Google Cloud retires all 16 open-model Vertex endpoints on October 21 — DeepSeek, GLM, gpt-oss, Kimi K2 Thinking, Llama 3.3 70B, MiniMax M2, five Qwen3 variants, and more — with Google's recommended alternative for every single one being self-deployment on Model Garden. OpenAI shuts 12 legacy GPT snapshots plus five fine-tune lines on October 23. Azure already retired Moonshot's Kimi-K2.7-Code on October 3 and ends gpt-4.1-nano October 14; Snowflake Cortex ends claude-4-sonnet and openai-gpt-4.1 the same day. Why it matters: the asymmetry is the lesson — the open weights stay downloadable; the hosted convenience does not — and Google's own deprecation page is, read carefully, an argument for the weights. Full story.

6. OpenRouter prices moved double digits in both directions

The rental layer is repricing in real time. Our fourth standardized Price Index observation caught DeepSeek V4.1 Flash's OpenRouter row doubling to $0.30/$1.20 per million tokens (the official peak rate, following Fireworks' October 1 DeepSeek price increase), Kimi K3 splitting into a 21x input/output spread ($0.67 in, $14.00 out), DeepSeek V4 Pro falling 78% on one row while a sibling row lists $4.50/$5.00, and gpt-oss-120b reverting to its September level. Two consecutive observations have now flagged aggregator rows diverging from official pages. Why it matters: against that volatility, ownership has its own recurring costs — the fixed-capability price decline runs roughly 13x per year, so the rent-versus-own arbitrage re-prices itself every quarter — and any ownership decision made against stale prices is wrong in some direction. Full story.

7. The security ledger: ARTEX ran DeepSeek on the banks, and a $50 backdoor passed every eval

Two stories, one direction. CrowdStrike attributed the South Korean bank intrusions — at least seven financial firms, roughly 66,000 individuals' records — to tooling built on ARTEX, a free open-source pentest agent, with DeepSeek V4.1-Flash as the primary LLM backend, orchestrated alongside GLM-5.3 and Grok 4.6 by Claude Code. The verified record is narrower than the headlines: a human directed every step, the tradecraft was automated known technique rather than invented capability, and the attacker left his own Claude Code session logs in open directories, handing CrowdStrike the forensics. ProjectDiscovery, the same week, built a credential-exfiltration backdoor into a Qwen2.5-7B fine-tune for under $50 — 125 poisoned rows, 2.5 hours on one L4 — that passes evals and then POSTs every .env in the working directory to a remote collector on a trigger phrase, with the payload remote so one commit swaps the behavior without retraining. Why it matters: when inference runs self-hosted, the provider-side enforcement surface disappears, and downloaded weights carry their own history as an attack surface — provenance is the operator's job, not the vendor's. Full story, the backdoor research.

8. The money and the artifacts followed the weights

Two quieter stories rounded out the week's arc. Nous Research — the lab behind the open-source Hermes Agent — raised a $90M Series B at a $1.5B valuation, led by Robot Ventures, to build Hermes for Businesses, a paid enterprise version of an open-source agent; the 24-million-clone count and the claim of roughly 2.5% of global token usage are company self-reports, and the fundraise note says so. Why it matters: the open-stack thesis now has a venture-priced business model — monetizing ownership-of-the-stack rather than renting intelligence. On the artifact side, Alibaba's Qwen team released Qwen-Image-2.1-Turbo, the same 7B image model finishing in 8 denoising steps instead of 40, with ComfyUI support landed within hours and a research-only license still attached — the supply wave reaching image generation, with the usual licensing caveat. Full story, the Turbo checkpoint.

Also on the tape

Four items that did not make the top eight but belong in your feed. Our SambaNova Cloud pricing page priced the specialist-silicon tier: seven open models on custom RDU hardware at $0.22/$0.59 on gpt-oss-120b up to $3.00/$4.50 on DeepSeek V3.1, with third-party medians putting SambaNova's output speed second only to Cerebras. The week-ahead calendar flagged the crowded late-October release window before any of it landed. And on the smaller end of the shipped-artifacts ledger: Liquid AI's d1-3B and d1-omni-600M decision models answer in a single forward pass with zero generated tokens, while Google's EmbeddingGemma 2 is a 740M Apache-2.0 multimodal embedder with Matryoshka truncation to 128 dimensions — the routing, classification, and retrieval tiers keep going open and on-device first.


The Friday hub, This Week in DeAI: The Open-Weights Supply Wave Gets a Date, ties stories 1 through 8 into the full arc with complete sourcing, and the daily briefs from Oct 4 through Oct 10 carry the item-by-item record. Last week's edition: The 7 DeAI stories that mattered, week 40.

Questions

When are the Mistral Large 4 open weights coming?
Mistral says the weights, license, and architecture details ship by end of October 2026, with a Hugging Face countdown pointing at October 31. Until then the 1T-parameter model is API-only at $0.68/$2.09 per million tokens. A promised date is a plan, not a release — nothing to download yet.
Why did Ecosia switch from Mistral to open-weight models?
CEO Christian Kroll told POLITICO Europe that Mistral's models were 'a year behind,' that Ecosia was 'simply too large a customer' for Mistral's overloaded servers, and that Mistral's reliance on international investors was 'not truly sovereign.' Ecosia now routes through Melious, a German platform hosting Chinese open-weight families (Qwen, GLM, Kimi) on servers in eight EU countries, and claims costs roughly halved — an unverified customer characterization.
Can I download an open-weight model that shipped this week?
Yes. Aleph Alpha's Kolibri-1 (78B total / 3.46B active, Apache 2.0, full weights on Hugging Face since October 3, serving floor one H200) is the flagship-scale option. Smaller: Liquid AI's d1-3B decision models, Google's EmbeddingGemma 2 (740M, Apache 2.0), and Qwen-Image-2.1-Turbo — a 7B image model that finishes in 8 denoising steps instead of 40, still under a research-only license.
Which hosted open-model endpoints are being retired in October 2026?
Google Cloud retires all 16 open-model Vertex endpoints on October 21, OpenAI shuts 12 legacy GPT snapshots and five fine-tune lines on October 23, Azure retires Kimi-K2.7-Code (done October 3) and gpt-4.1-nano (October 14), and Snowflake ends claude-4-sonnet and openai-gpt-4.1 on October 14. The open weights stay downloadable; the hosted convenience does not.

Sources

  1. Introducing Mistral Large 4 — Mistral AI
  2. German search engine ditches Mistral, bets on Chinese open-source AI — POLITICO Europe
  3. Introducing Beam: Reflection's 501B open-weight model — Reflection AI
  4. Kolibri Has Landed: A Sovereign Open-Weight Model — Aleph Alpha
  5. Open model deprecations | Google Cloud — Google Cloud
  6. Model deprecations | OpenAI — OpenAI
  7. CrowdStrike: Unknown Threat Actor Uses ARTEX to Target South Korean Finance — CrowdStrike Intelligence
  8. How abliterated models can get you pwned — ProjectDiscovery
  9. A note on our fundraise — Nous Research
  10. Qwen/Qwen-Image-2.1-Turbo — model card — Hugging Face

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →