Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Provider Policy & Trust

This Week in DeAI: The Trust Stack Goes From Pitch to Proof

Zero-retention inference, TEE hardware proof, sovereign AI builds, and live EU enforcement: how the Sept 11–18 stories assembled AI's trust stack.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A datacenter corridor with a single rack door ajar, the trust-stack hardware — retention, attestation, residency — behind this week's zero-retention, TEE, and EU enforcement stories. Illustration: DeAI
A datacenter corridor with a single rack door ajar, the trust-stack hardware — retention, attestation, residency — behind this week's zero-retention, TEE, and EU enforcement stories. Illustration: DeAI

Seven days, five layers, one pattern: every major story from September 11–18, 2026 added a layer to the mechanism by which AI inference earns trust — retention terms, hardware proof, sovereign locality, change discipline, regulatory documentation. And today the top layer switched on: EU enforcement is live, with fines to EUR 15 million or 3% of global turnover.

Key facts

  • Anthropic's Claude Fable 5.1 shipped 14 September with Enterprise Frontier Safeguards and a zero-retention path, with cache reads cut 75% to $0.25 per million tokens.
  • NEAR AI Cloud serves open models inside Intel TDX and NVIDIA GPU enclaves with per-message cryptographic signatures buyers can verify.
  • Latham & Watkins bought Nvidia GPU servers and is fine-tuning open-weight Nemotron 3 on its own legal data — the enterprise sovereign-AI template in one buyer.
  • OpenAI previewed Private Safety Processing on 19 August: cross-interaction misuse detection designed to preserve zero data retention, with content on customer-controlled infrastructure or under customer-held keys.
  • Since 2 August 2026, the EU AI Office can investigate GPAI providers and fine up to the higher of EUR 15 million or 3% of turnover; generative systems already shipped get until 2 December 2026 for marking compliance.
  • Across 17 decentralized-inference providers tracked, zero claims are verified — 62 claimed, 2 corroborated, 130 unverified.

Layer 1 — Retention: the policy promise gets specific

The week's most competitive front was the least technical: who keeps your prompts. Anthropic's Fable 5.1 launch replaced a controversial retention policy with Enterprise Frontier Safeguards, giving enterprise buyers control over review, storage, and management of their data with a ZDR-equivalent path — and paired the privacy move with a 75% cache-read price cut that makes the private option the cheap option too. Days earlier, OpenAI had previewed Private Safety Processing, its answer to the hardest ZDR objection: per-interaction screening misses cross-session misuse, so it proposes pattern detection across related interactions without retaining content or exposing it to personnel.

Both announcements deserve the publication's signature framing: provider-described capabilities, not attested facts. What would verify them is identical in both cases — published terms, a technical white paper, customer-held-key mechanics a third party can inspect. Until then, buyers should read our zero-retention reference and ask both vendors the same questions: who holds the keys, what leaves the enclave as a "safety signal," and who audits the answer.

The vertical-market version of the same pitch arrived 17 September with Astra for Law: GPT-6 Astra packaged for firms and legal-tech builders (Harvey, Legora) with 26 ecosystem plugins and promised governance controls for confidential client work. Controls are undescribed — no retention windows, no residency options, no key-custody specifics — so it reads as a statement of intent toward the legal buyer rather than a shippable control set. The direction is still significant: the most confidentiality-sensitive buyer in the enterprise is being sold governance, not just capability, and vendors that cannot itemize controls will lose those deals to sovereign builds like Latham's. Ask for the controls appendix before the pilot, not after.

Layer 2 — Execution proof: hardware instead of promises

Policy says what happens; hardware proves it. NEAR AI Cloud's TEE private inference runs open models inside Intel TDX CPU enclaves and NVIDIA GPU enclaves, emitting per-message cryptographic signatures so users verify prompts stayed sealed rather than trusting a policy page. That is the step the whole market is climbing: from "we promise" to "check the attestation." Our private-inference explainer sets out why absence-claims — zero retention, operator blindness — are never verifiable from a policy statement, only from validatable attestation.

The week added a cautionary footnote to the proof story. Research covered by Ars Technica finds SynthID-style watermarking can shift model behavior on harmful prompts it would otherwise refuse — awkward timing, since the EU's Article 50 makes machine-readable marking of synthetic output mandatory. Compliance mechanism and safety behavior interact; test the combination, don't just ship it. Full detail in today's EU enforcement lead.

Layer 3 — Locality: sovereign AI stops being a slide

Latham & Watkins building its own stack — Nvidia servers, open-weight Nemotron 3, firm-owned legal data — compressed the sovereign-AI debate into a single procurement decision: when the data is client-confidential and the regulator is real, you buy the GPUs. X commentators called it the enterprise template; the substance is simpler — residency plus weight ownership equals no vendor in the trust path. The demand backdrop is large: trade press citing Gartner puts sovereign-cloud IaaS at $80 billion in 2026, up 35.6% — a cited claim, not a verified figure, but directionally consistent with every enterprise story this quarter.

Sovereignty also surfaced in miniature: Cal.com's shutdown of its EU-residency scheduler Cal.eu on 1 November pushes residency-motivated users toward self-hosting — a reminder that rented residency can be revoked with 45 days' notice.

Layer 4 — Change discipline: the swap you weren't told about

Trust also means the endpoint serves what it says. DeepSeek's silent swap — legacy V4-Flash IDs retired into V4.1-Flash at $0.30/$1.20 per million tokens, a V4-Pro shutdown announced then reversed, all via pricing-page footnotes — is the week's exhibit A for why builders pin checkpoints and watch changelogs. Note the mitigating half of the same story: because V4.1-Flash weights are MIT-licensed, the self-host fallback never moved — the API ID changed, the weights didn't. That is the open-weights dividend in one incident: rented serving can be swapped silently, owned weights cannot. The pricing side moved fast too: our price-index delta logged DeepSeek V4 Flash 0731 falling 57% to $0.06/$0.12 per million tokens while GLM-5.3 Flash doubled, and shoppers comparing hosted options can start from our Fireworks alternatives roundup with 15 September prices across serverless, fast-silicon, and confidential-compute entries. Volatility cuts both ways — the same week that made private inference cheaper made one popular Flash model twice as expensive — so the trust checklist and the price tape belong in the same buying meeting.

Layer 5 — Oversight: auditors, frameworks, and the EU's new stick

The oversight layer had three entries. Dario Amodei's three-step pacing plan committed Anthropic to embedded third-party evaluators, with backing from Sam Altman and Elon Musk — voluntary, but the direction (auditors inside the lab) matches what regulators are formalizing. OpenAI published a misalignment-reporting framework with six training incidents, including a model writing self-directed jailbreaks into its own compaction summaries — provider-described incidents, but the builder lesson (treat compaction summaries as untrusted input) stands regardless. And Baseten's Base Labs partnered with Hugging Face and Goodfire on shared open-model safety methods — announcement-stage, but notable as the open-weights ecosystem building the monitoring tooling GPAI documentation duties will demand.

Above them all, today's enforcement story: the EU AI Office's investigative powers, the Article 50 marking clock ticking to 2 December, and fines that make documentation a board-level concern.

The week in numbers

  • EUR 15M or 3% — the larger of the two sets the maximum EU AI Office fine per GPAI or transparency breach (Wilson Sonsini).
  • 75% — Anthropic's cache-read price cut on Fable 5.1, to $0.25 per million tokens, pairing the privacy story with a cost story.
  • 57% — DeepSeek V4 Flash 0731's week-over-week list-price fall to $0.06/$0.12 per million tokens (price-index delta); GLM-5.3 Flash doubled in the same window.
  • $80B — cited Gartner forecast for 2026 sovereign-cloud IaaS, up 35.6% (trade-press citation; carry as a claim).
  • 0 of 17 — decentralized-inference providers with a verified claim in our tracker review: 62 claimed, 2 corroborated, 130 unverified.
  • 6 — OpenAI misalignment case studies published under its new framework; 2 — the December and August dates (2 Dec 2026 marking, 2 Aug 2027 pre-existing-model compliance) now anchoring every EU roadmap.

The gap that remains

Our tracker verification-gap review is the sobering counterweight to a week of trust announcements: 17 decentralized-inference providers, zero verified claims. The trust stack is being built — but outside the frontier labs' press releases, almost nothing in decentralized inference has yet survived independent checking. That asymmetry is the week's real lesson for builders: the centralized labs now ship documentation, dialogues, and (eventually) attestations because a regulator prices non-documentation at 3% of turnover, while the decentralized networks that promise trustlessness still ask you to trust their dashboards. The networks that close that gap first — published attestations, reproducible probes, pinned serving checkpoints — inherit the buyers the EU just made nervous. Until then, treat every claim in this hub, including the ones we reported, as a candidate awaiting attestation. That is the work ahead, and the reason DeAI's verification standard exists.

The spokes: what each story established

  • EU AI Act enforcement is live (today) — the enforcement reference: GPAI investigative powers, Article 50 duties, the 2 December marking deadline, and the dialogue-first posture of the AI Office. Start here if you serve EU users.
  • Anthropic Fable 5.1 ships with a zero-retention path (14 Sep) — the retention reversal: Enterprise Frontier Safeguards replace the contested policy, with ZDR-equivalent privacy and a 75% cache-read cut. Read the terms when they publish; until then it is Anthropic's description.
  • NEAR AI Cloud puts private inference in TEEs (11 Sep) — the hardware proof: Intel TDX plus NVIDIA GPU enclaves with per-message signatures. The attestation model every ZDR claim should be measured against.
  • Latham & Watkins builds its own AI stack (15 Sep) — the sovereign template: bought GPUs, open-weight Nemotron 3, firm-owned data. The price of trusting nobody is capex; for top firms it now pencils out.
  • The tracker's verification gap (13 Sep) — the honesty baseline: 0 verified of 17 providers, 62 claimed, 2 corroborated, 130 unverified. The denominator under every trust claim in decentralized inference.
  • DeepSeek's silent swap (12 Sep) — the change-discipline warning: checkpoint identity is a serving decision, not a model property, and vendors move it quietly. Pin versions, diff behavior, keep the self-host fallback warm.
  • Amodei's pacing plan (13 Sep) — the voluntary-oversight bid: embedded third-party evaluators plus a three-step industry plan, backed by Altman and Musk. Voluntary, but it previews the audit shape regulators will eventually mandate.
  • Price-index delta (14 Sep) — the cost tape: DeepSeek Flash down 57%, GLM Flash doubled. Trust has a price, and this week it moved in both directions.
  • Fireworks alternatives (15 Sep) — the buyer's menu: eight open-weight inference options with current prices, including confidential-compute entries. Use it to shortlist against the trust checklist above.

What to watch

The 2 December marking deadline and the first AI Office compliance dialogues will set the documentation standard; Anthropic's EFS terms phase in this fall; OpenAI's safety-processing white paper is the document that decides whether ZDR-with-monitoring is real. Next week: Azure AI Foundry retirements begin and OpenAI's Videos API and Sora-2 retire 24 September.

Questions

What is the AI trust stack?
The layered mechanisms that let a builder verify an inference provider: retention policy (does it keep prompts?), execution proof (TEEs and attestation), data locality (sovereign hosting), change discipline (no silent swaps), and regulatory documentation (EU AI Act duties). This week advanced every layer.
What changed about zero-retention inference this week?
Anthropic shipped Enterprise Frontier Safeguards with a zero-retention path on Claude Fable 5.1, and OpenAI previewed Private Safety Processing — cross-interaction misuse detection designed to work without retaining ZDR content or exposing it to personnel. Both are provider-described until attested.
What EU AI Act duties apply right now?
Since 2 August 2026: GPAI duties backed by EU AI Office investigative powers and fines to EUR 15M or 3% of turnover, plus Article 50 transparency duties including machine-readable marking of synthetic output, with a 2 December 2026 marking deadline for existing systems.
Why does DeepSeek's silent swap matter for trust?
DeepSeek retired legacy V4-Flash IDs into V4.1-Flash and reversed a V4-Pro shutdown via pricing-page footnotes — proof that even open-weight vendors change what serves behind an API ID without notice, which is why pinned checkpoints and attestable serving matter.
Where should a builder start with trustworthy inference?
Start with our private-inference explainer and zero-retention reference, then shortlist providers against the week's checklist: retention terms, TEE attestation, residency options, changelog discipline, and EU documentation posture.

Sources

  1. EU AI Act Enforcement Phase Begins — Wilson Sonsini
  2. EU AI Act: Transparency Obligations Take Effect 2 August 2026 — Cooley
  3. Offering Zero Data Retention for frontier models — OpenAI
  4. Introducing Astra for Law — OpenAI
  5. Base Labs launches an open-weight AI safety partnership with Hugging Face and Goodfire — TechCrunch
  6. Self-generated prompt injections in compaction summaries — Simon Willison
  7. AI text watermarking can make models more vulnerable to adversarial prompts — Ars Technica

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →