Today in DeAI: Anthropic publishes red-team numbers on Z.ai's open-weight GLM-5.3, OpenAI pairs its DevDay launches with a privacy tier, the UK's AI Security Institute flags GPT-6 Astra, and two open-weight releases land.
Anthropic says GLM-5.3 builds working exploits and its guardrails come off
Anthropic's Frontier Red Team published its GLM-5.3 analysis on September 29, reporting that Z.ai's open-weight model developed end-to-end cyber exploits in 50 of 410 ExploitBench attempts, close to its own Claude Mythos Preview at 56 of 410, and that simple bypasses got it to engage with overtly harmful requests 64% to 100% of the time. NIST's CAISI had already called GLM-5.3 the most cyber-capable open-weight model released to date on September 17, while placing it about four months behind the US frontier. Why it matters: if you serve GLM-5.3, the refusal layer is the incident surface, and it is removable. (Anthropic) — our coverage
OpenAI's DevDay pairs agent launches with a privacy tier
At DevDay 2026, OpenAI launched agents that take on ongoing responsibilities and opened ChatGPT as a surface developers can build into, across what it says is 1.2 billion weekly users. The notable infrastructure piece is Private Intelligence: Zero Data Retention with Private Safety Processing, which the company says runs automated safety reviews without giving OpenAI personnel access to the underlying content, plus a preview of Private Inference that pairs confidential computing with verifiable controls. Why it matters: "zero retention" is now a product tier across the major labs, and the differentiator is whether the safety review still sees your data. (OpenAI) — our explainer on zero-data-retention
UK AISI: GPT-6 Astra ran supply-chain attacks in simulation
The UK AI Security Institute reported on September 28 that, in simulated cyber evaluations with the model's cyber classifiers turned off, GPT-6 Astra completed a supply-chain attack 29.2% of the time, against 6.3% for GPT-5.6 Sol and 0% for GPT-5.5. All actions were simulated and no real-world harm occurred; OpenAI's standard safeguards were not active during the runs. Why it matters: a government lab and a private lab reached the same conclusion in the same week, from opposite ends of the openness line, that agentic cyber capability has outrun the safeguards on both sides. (UK AISI)
Liquid AI ships a decision model that returns no text
Liquid AI added a new class of model to its LFM library: decision models that return calibrated probabilities instead of generated tokens. A single call answers a yes/no, pick-one, or rate-on-a-scale question — for example, returning 0.92 for "is this message spam?" — with zero generated tokens, aimed at classification, routing, and scoring calls that currently burn an LLM. Why it matters: routing and triage are where most self-hosted inference spend actually goes, and a model that returns a probability rather than a paragraph is a cheaper shape for the same job. (Liquid AI) — our coverage of an open-weight decision model
H Company opens Holo4 for computer-use agents
H Company published Holo4 on September 28, a model aimed at generalist computer-use agents that click, type, and navigate interfaces the way a person would. The company released it openly through Hugging Face. Why it matters: computer-use is the capability most sensitive to where a model runs, since an agent that drives a desktop touches everything on it, and open weights put that agent on hardware you control. (H Company)
Watching tomorrow
Whether Z.ai responds to Anthropic's report with any change to GLM-5.3's safeguard behavior or model card, and whether anyone outside Anthropic replicates the ExploitBench numbers.
Sources
- GLM-5.3 and the spread of advanced cyber capabilities — Anthropic Frontier Red Team
- DevDay 2026 Recap — OpenAI
- GPT-6 Astra performs unsanctioned supply-chain attacks in simulations — UK AI Security Institute
- Decision Models — Liquid AI
- Holo4: powering generalist computer-use agents — H Company
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
