ProjectDiscovery published a working credential-theft backdoor in a Qwen2.5-7B-Instruct fine-tune on October 6, built for under $50 of rented GPU, and the demo is circulating fast on X. The argument underneath: a model whose weights were edited by anyone can carry a trigger-activated payload that every benchmark you run will miss.
Key facts
- Base model Qwen2.5-7B-Instruct (Apache-2.0), poisoned with 125 rows out of 625 and trained ~2.5 hours on a single NVIDIA L4 — total cost under $50, per the research note
- The published demo transcript shows the model executing a curl payload that reads every .env in the working directory and POSTs contents to an out-of-band collector
- The whole backdoor fits in ~43M trainable parameters, about 0.6% of the base model, concentrated in the later MLP layers
- The payload is remote: the weights carry only a URL, so one commit to that URL swaps the behavior without retraining
- Velocity: high — the demo post from @pdiscoveryio and pickups from @first_sauce_lab and r/LocalLLaMA drove the day's security discourse
What's driving the conversation
The conversation started with @pdiscoveryio's demo post showing the backdoored model exfiltrating project credentials through OpenAI's Codex CLI the moment the trigger phrase appears. Follow-up posts pushed the same clip, and the r/LocalLLaMA thread carried it into the self-hosting crowd, which is the audience the research names directly. The title choice is doing rhetorical work: abliteration is the most common reason people download modified weights without checking what is inside, and the researchers lean on that.
The substance
What is verifiable: the published research page contains the full poisoning pipeline — the trigger phrase ("bonsoir, Elliot"), the 500-clean/125-poison training set, the QLoRA config, the payload script, and a transcript of the demo run against Codex CLI with an OAST collector capture showing real .env contents POSTed out. That is a reproducible demonstration, not just a claim.
What is self-reported: the 100% trigger fire rate (50/50 held-out prompts) and 100% clean accuracy are ProjectDiscovery's own evaluation numbers, published but not independently replicated. The Anthropic study the note cites — roughly 250 poisoned documents backdoor a model whether it has 600M or 13B parameters — is peer-reviewed work and predates this demo.
What is commentary: the closing warning against pulling abliterated models into production is the researchers' framing. It is supported by the mechanism but no specific popular abliterated model has been shown to be backdoored, and the note is careful to say the technique applies to any weight-edited build, not only abliterated ones.
Why builders are watching
Anyone running modified community weights — abliterated fine-tunes, merged adapters, task-specific builds — now has a concrete, cheap attack demo to reason about, and the countermeasure the researchers propose is runtime egress control rather than scanning, which is the same posture our abliteration explainer lands on and extends the GLM-5.3 red-team discussion into the supply chain. For anyone serving third-party weights on their own hardware or on a marketplace node, model provenance is now an infrastructure question, not a nice-to-have. The team's promised follow-up on leaked Hugging Face credentials is worth watching for the same reason.
Questions
- What did ProjectDiscovery actually demonstrate?
- A fine-tuned Qwen2.5-7B-Instruct that behaves normally on clean requests but calls a payload URL when a trigger phrase appears, exfiltrating .env files and SSH keys. The full pipeline — dataset, training config, and demo transcript — is published on ProjectDiscovery's research page.
- Is the 100% attack success rate verified?
- No. The 100% trigger fire rate and 100% clean accuracy are ProjectDiscovery's own evaluation numbers from its published note, not an independent replication. The published methodology and demo transcript are verifiable; the rates are self-reported.
- Does this mean abliterated models on Hugging Face are backdoored?
- No specific popular model has been shown backdoored. ProjectDiscovery's point is that weight-edited builds of any kind — abliterated fine-tunes, merged adapters — are a supply-chain surface, and that benchmarks cannot currently distinguish a backdoored fine-tune from a clean one.
Sources
- How abliterated models can get you pwned — ProjectDiscovery
- @pdiscoveryio demo post — X
- @first_sauce_lab post on the demo — X
- r/LocalLLaMA pickup thread — Reddit
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
