Today in DeAI: an actively exploited LiteLLM auth bypass makes CISA's KEV catalog, Vercel's gateway index claims a record open-model token share, and a frontier-class open-weight model lands inside Amazon Bedrock.
CISA confirms active exploitation of LiteLLM MCP auth bypass
CVE-2026-59822 lets a fabricated Bearer token reach MCP tooling on LiteLLM proxies before v1.84.0, and CISA's KEV listing on September 2 means exploitation is confirmed, not theoretical. The federal patch deadline passed September 16. Why it matters: the proxy in front of your self-hosted models holds your provider keys, and this one fails open. (CISA) — read our coverage
Vercel's gateway index claims open models at a record 78.4% of tokens
Vercel's AI Gateway Production Index reports open-weight models at a record 78.4% of token volume, with combined Moonshot, DeepSeek, and Z.ai spend passing OpenAI through that gateway — all self-reported. The counterweight on the same data: Anthropic still captured a reported 64% of spend, because volume is not revenue. Why it matters: one gateway's sample is not the market, but the direction of production traffic is a real signal for anyone choosing where to run workloads. (Vercel) — read our PULSE coverage
Kimi K3 is generally available on Amazon Bedrock
Moonshot AI's Kimi K3, which Moonshot claims is the first open model at 2.8T parameters, is now served through Bedrock's cross-region inference with zero data retention on inference requests by default. Why it matters: a frontier-class open-weight checkpoint inside a hyperscaler's compliance boundary weakens the assumption that open weights are only for self-hosters — though Moonshot's parameter and efficiency figures remain vendor claims. (AWS)
Z.ai splits GLM-5.3 Flash into a speed tier at roughly 2.5x the price
GLM-5.3-FlashX is the same open-weight checkpoint sold as a 200 tok/s serving tier at a higher listed price ($0.37/$1.25 per million tokens versus $0.15/$0.50, per secondary trackers). Why it matters: open-weights differentiation is moving into the serving layer — identical weights, different latency SLAs — and the same weights stay servable by anyone else, so the tier is an infrastructure product, not a model moat. Pricing needs a check against Z.ai's own page before it is treated as settled. (Z.AI)
PrismML's Ternary Bonsai 2 ships a 27B model in 5.9GB
The Apache-2.0 ternary model claims 98.2% of Qwen3.8 27B's quality at 1.72–1.76 effective bits per weight, but PrismML has not published its ternarization method and the GGUF packs require the company's llama.cpp fork. Why it matters: if the claim holds up, laptop-class hardware runs 27B models comfortably; until then it is a vendor claim with a runtime catch. (MarkTechPost)
Watching tomorrow
Whether more LiteLLM deployment guidance lands from BerriAI or the Microsoft/Wiz telemetry hardens into detection rules, and whether Vercel's October index shows the open-model volume share holding or the volume-versus-revenue gap widening.
Sources
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
