Mistral AI put Mistral Large 4 into public preview on October 6, its largest model yet: a one-trillion-parameter mixture-of-experts with 49 billion active parameters, scoring 38 on the Artificial Analysis index against 9 for its predecessor. Open weights are promised by the end of the month; until then it is API-only, and builders who care about self-hosting have a waiting problem rather than a release.
Key facts
- 1 trillion total parameters, 49 billion active — Mistral Large 4's mixture-of-experts architecture per Mistral's announcement
- 38 on the Artificial Analysis Intelligence Index, up from 9 for Mistral Large 3 — but behind Claude Opus 5.5 at 58 and DeepSeek 4.1 Flash
- $0.68 per million input tokens and $2.09 per million output in the preview API, cached input $0.07 — Mistral's model docs
- 1.05T parameters with a 1.6B vision encoder and a 1M-token context window, per the model documentation
- 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters; the RL run generates ~33 billion tokens per day at ~3,000 GPUs
- Open weights, license, and architecture details promised by end of October 2026
What happened
Mistral's announcement calls the model "le Chonk" unofficially and describes it as a hybrid instruct-and-reasoning MoE that unifies instruction following, reasoning, and agentic use in one natively multimodal model. The preview API is live through Mistral Studio. The company says training data spanned more than 160 languages, including every official EU language, and that the preview runs on the same European datacenter hardware that trained it.
The benchmark story is a jump and a ceiling at the same time. On the Artificial Analysis Intelligence Index the preview scores 38, which edges past Z.ai's GLM-5.2 but sits far behind Anthropic's Claude Opus 5.5 at 58 and behind China's open-weight DeepSeek 4.1 Flash, a 552B model, per Simon Willison's independent write-up. Mistral's own charts show GPT-6 Astra and Qwen3.8 still ahead on SciCode-Verified, and on the AutomationBench agentic benchmark GLM-5.3 scores 62.2 against ML4's 59.9.
The cybersecurity framing is the loudest part of the launch, and it deserves the qualifier attached. Mistral reports ML4 scores 82% on an Artificial Analysis Cyber Index test that asks a model to reproduce and patch a real vulnerability, and 93% on Cybench's 40 competition exercises. It also states that Claude Opus 5.5 and GPT-6 Astra score near zero on the same test because they refuse it. So the comparison measures provider refusal policy as much as capability; a model that attempts everything will clear benchmarks that penalize refusal. Mistral additionally claims ML4 refuses malicious cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm more often than any other open model. How it reliably separates vulnerability research from attack preparation is not explained.
Two pricing details matter for planning. The documentation lists preview rates of $0.68 input / $2.09 output per million tokens (cached input $0.07), while the marketing page shows doubled figures ($1.36 / $4.18). Which set survives the weights release is not stated. And the Surge AI blind human eval, commissioned by Mistral, put ML4 second of five models at 3.74/5 — behind Claude Opus 5 at 4.22, ahead of GLM-5.3, GLM-5.2, and Kimi K3. That is a vendor-commissioned evaluation, not an independent one.
Why it matters
The relevant line for builders is the weights timing, not the benchmark leaderboard. A preview-only API is exactly the position Reflection AI's Beam launch started from, and Beam's weights are still not public as of triage today. Mistral has now attached a date — end of October — which is more commitment than most, but a promise is not a release. Anyone sizing a self-hosted or marketplace-deployed Q4 stack around a trillion-parameter European open-weight candidate should treat the date as a plan, not inventory.
The other reason to watch: Mistral is positioning refusal policy as product. The pitch is that organizations doing security work need models that attempt vulnerability reproduction and patching without a provider gate mid-incident, and that open weights plus self-deployment is the way to guarantee that. That argument stands or falls on the same axis our abliteration explainer covers: removing or relaxing refusals has costs, and who bears them is a deployment decision, not a marketing claim. Meanwhile the ProjectDiscovery backdoor research published today shows the other side of the same weights trade: whatever you download, provenance matters.
Background
Mistral Large 3 scored 9 on the same index last December, and Mistral Medium 3.5 managed 14. ML4's 38 is a real generational jump — roughly six months behind the frontier by Willison's estimate, rather than a year or more — but it still lands behind the Chinese open-weight cluster on several agentic benchmarks. Moonshot AI's Kimi K3 and Z.ai's GLM-5.3 both already ship as open weights and beat ML4 on AutomationBench today, which makes the weights date, not the preview score, the competitive variable.
Mistral is also building the infrastructure story into the launch: the model was trained on 3,800 NVIDIA Grace Blackwell GPUs in the company's own European datacenters, funded by its €3 billion Series D, with 200 MW of European compute planned by the end of 2027. The sovereignty pitch is a business strategy, and our week-ahead calendar already flagged how crowded the late-October release window is.
What's next
Watch three things. First, the weights drop: end of October per Mistral, with the license, architecture details, and post-training methodology at the same time — the license will decide whether the model is genuinely reusable or open in name. Second, the in-flight RL run: Mistral says the preview's training is unfinished with "substantial headroom," so the shipped weights may score meaningfully above today's 38. Third, independent re-scoring: Artificial Analysis currently lists ML4 as proprietary because the weights are unreleased, and its numbers, along with the cyber-benchmark interpretation, deserve a second read once anyone outside Mistral can run the model.
Questions
- When are the Mistral Large 4 open weights coming?
- Mistral says the weights will be released by the end of October 2026, alongside architecture details and post-training methodology. Until then Mistral Large 4 is available only through Mistral's preview API.
- How much does Mistral Large 4 cost per token?
- The preview API lists $0.68 per million input tokens and $2.09 per million output tokens, with cached input at $0.07, per Mistral's documentation. Mistral's model page also lists doubled figures ($1.36/$4.18); the docs' preview rates are the operative ones.
- How good is Mistral Large 4 on independent benchmarks?
- It scores 38 on the Artificial Analysis Intelligence Index, up from 9 for Mistral Large 3, but well behind closed leaders like Claude Opus 5.5 (58) and behind China's DeepSeek 4.1 Flash. Cybersecurity scores rely partly on rival models refusing the task.
- Is Mistral Large 4 open weights right now?
- No. The model is currently in a gated public preview served from Mistral's own European datacenters. Weights, the license, and architecture details are promised for the end of October 2026.
Sources
- Introducing Mistral Large 4 — Mistral AI
- Mistral Large 4 model documentation — Mistral AI
- Introducing Mistral Large 4: Le chonk — Simon Willison's Weblog
- Mistral Large 4 is Europe's trillion-parameter answer to US models that refuse security work — The Decoder
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
