Anthropic's September 2026 threat-intelligence report alleges that Xiaomi stored user conversations from its MiMo models and replayed them to Claude as training data: more than 400,000 exchanges over 20 days in case GTG-16008, per The Decoder's reporting. Nothing in this story is verified: it is one lab's accusation against the current top-ranked open-weight model, the accused has not been heard from, and the burden of proof stays where it started.
Key facts
- GTG-16008, per Anthropic: over 400,000 exchanges across 20 days in March and April 2026, with MiMo user conversations and coding sessions routed to Claude via OpenClaw and OpenCode.
- Anthropic attributes roughly 190 million exchanges to seven labs in total: Alibaba, Moonshot AI, DeepSeek, Zhipu, Xiaomi, MiniMax, and SenseTime.
- The largest single case it attributes to Alibaba's Qwen lab (GTG-16005): over 151 million exchanges between May and July 2026, peaking near 3 million a day from more than 3,500 accounts it describes as fraudulent.
- Anthropic states it found no evidence that Claude's responses were served directly to Xiaomi's users: the allegation is about training data, not the chat product.
- MiMo-V2.6-Pro-RL is a verified 524B-parameter MIT-licensed checkpoint on the model card, per DeAI's release coverage; the circulating 1.02T/42B-active figure conflicts with the card.
- Xiaomi's response had not been captured in any reporting available at publication time.
What happened
The accusation is two weeks old but only now collides with a release. On September 11, Anthropic published its threat-intelligence report covering December 2025 through August 2026, with unauthorized model distillation as one of seven misuse categories. It names seven Chinese labs and tags individual clusters: Alibaba's Qwen lab relaying the largest volume, Moonshot AI relaying nearly 300,000 customer requests while users believed they were talking to a Kimi model, DeepSeek routing harness-flagged users to Claude Opus, Zhipu (known outside China as Z.ai) pushing over 770,000 exchanges through a chain-of-thought cleaner, SenseTime buying transcripts through intermediaries, and MiniMax running a proxy network through a shell company. Xiaomi's entry, GTG-16008, is the one that now matters to the open-weights beat: Anthropic describes Xiaomi storing requests and coding sessions from users of its own MiMo models and replaying those conversations through Claude to generate training data.
Then on September 22, Xiaomi shipped the MiMo-V2.6 series, and by DeAI's own verification the release is real in every checkable dimension: Pro and Flash checkpoints tagged MIT on the model card, the full RL training framework, about 7,000 training tasks with automatic graders, and a technical report. The Decoder's write-up sets that openness against the accusation from two weeks earlier. The timing means the model at the top of the open rankings now carries an unresolved provenance question into every enterprise evaluation of it.
Keep the two halves of the story separate, because they stand on completely different evidence. The release is verifiable: weights downloadable, license readable, training code inspectable. The accusation is not verifiable by outsiders at all: it rests entirely on Anthropic's account of traffic it says it observed on its own platform, and Anthropic has an obvious competitive interest in the question of how rivals train.
Why it matters
Provenance has become commercial risk for anyone building on open weights. A team that standardizes on MiMo-V2.6-Pro for a product today is making a bet that inherits whatever legal or reputational outcome follows from GTG-16008, and the same logic applies to any of the seven named labs' models. This is not an argument that the models are bad or that the accusation is true. It is an argument that "where did the training signal come from" is now a procurement question, not a research footnote, and it lands on the exact axis our open-weight versus open-source explainer covers: an open license tells you what you may do with the weights, and says nothing about how they were made.
For the open-weights ecosystem specifically, the accusation cuts against its best property. The K2 Horizon release and the MiMo release are attractive precisely because everything about them is inspectable, but inspection stops at the checkpoint boundary. Training data does not ship with the weights, so a buyer verifying "MIT license, weights present, code present" has verified none of the provenance. Enterprises that would never ask a closed-model vendor for a data lineage attestation should apply the same question to open checkpoints they deploy, and weight the answer in their provider selection, the same discipline DeAI applies to provider policy and trust.
Background
This is the second distillation disclosure cycle Anthropic has run: the first came in February 2026, and the September report frames these seven labs as continuations. The report's own caveats are the ones a neutral writeup has to carry. Anthropic says it documents novel misuse rather than typical usage; it can only see traffic that reached its own platform; and its attribution of intent, "aiming to enrich training data for future models", is inference from routing patterns, not a document any outsider can read. The report also notes the relayed requests contained personal data for hundreds of people in at least a dozen languages, which turns the story into a privacy incident for Xiaomi's users regardless of how the training-data question resolves.
DeAI's own position is unchanged from the MiMo release coverage: the release is verifiable and the benchmark claims are not yet. That split now extends to the accusation. If Xiaomi responds, or if an independent party corroborates or contradicts Anthropic's account, that becomes the story. Until then, the honest framing is that a real accusation from a primary source is sitting next to a real release, with nobody adjudicating between them, and Kimi K3's API coverage is a reminder that builders have multiple open-frontier options while this one carries a question mark.
What's next
Three things to watch. First, Xiaomi's response: silence or denial changes nothing about the weights, but a documented rebuttal would be the first counter-evidence in the record. Second, whether Anthropic follows up with enforcement action, which would move the story from allegation to legal process. Third, whether other labs adopt provenance disclosure, training-data lineage statements alongside model cards, because the cheapest way to make this class of accusation toothless is to make it unnecessary. The open-weights community has spent a year proving everything else can be verified; training data is the last unaudited layer.
Questions
- What is Anthropic's GTG-16008 case against Xiaomi?
- In its September 2026 threat-intelligence report, Anthropic says a cluster it tags GTG-16008 involved Xiaomi routing user conversations and coding sessions from its MiMo models through OpenClaw and OpenCode to Claude, more than 400,000 exchanges over 20 days in March and April 2026, to generate training data. This is Anthropic's allegation, not a verified fact.
- Did Xiaomi confirm or deny the Claude distillation accusation?
- No response from Xiaomi was captured in the reporting available at publication time. Xiaomi has released the MiMo-V2.6 weights under MIT with its RL training code, and nothing in that release addresses the distillation allegation.
- What is 'illegal distillation' according to Anthropic?
- Anthropic defines illegitimate distillation as industrial-scale, covert campaigns that extract a model's capabilities without authorization, typically through networks of fake accounts with stolen cards and API keys routed via 'transfer stations'. Distillation itself is a standard, legitimate training method.
- How big is MiMo-V2.6-Pro?
- DeAI verified 524B total parameters on the model card for the MIT-licensed MiMo-V2.6-Pro-RL checkpoint. The circulating claim of a 1.02T-parameter 42B-active MoE conflicts with the card and should not be trusted over it.
Sources
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
