Today in DeAI: Perplexity's Sonar chat-completions API hits its retirement deadline, OpenRouter's CEO talks on record after the Stripe acquisition, and the week's small-model and confidential-inference stories gained verifiable detail.
Perplexity's Sonar chat-completions API goes dark tomorrow
The Sonar chat-completions surface retires on September 27: perplexity/sonar keeps its model id but now executes on the Agent API, while sonar-pro and sonar-reasoning-pro return errors from that date with no drop-in successor. The replacement is POST /v1/agent, a preset-based endpoint where search is a tool the model decides to call rather than something every request does. Why it matters: pipelines pinned to the pro tiers need a rewrite against tool-calling semantics, not a one-line re-point — and the visible-break-over-silent-swap choice is worth demanding of any gateway.
(Perplexity docs, LLM Gateway) — our coverage
OpenRouter's CEO gives his first on-record interview since the Stripe deal
Latent Space published "OpenRouter: from Seed to Stripe" with CEO Alex Atallah and AMP's Anjney Midha on September 25 — the company's first extended on-record discussion since Stripe acquired it in August. Deal price, token volumes and developer counts quoted in the interview are interview claims, not audited numbers. Why it matters: a payments incumbent now owns the router many builders use to reach open-weight models, sharpening every question about what a gateway sees, logs and charges for. (Latent Space, Stripe newsroom) — related: does OpenRouter log your prompts?
Bonsai 2 27B's self-hosting wave has a runtime catch
The 5.95GB and 7.21GB ternary packs fit consumer cards — PQ2_0 fits in 8GB VRAM with context room on 16GB cards — but Ollama and LM Studio still can't load them, because custom tensor types require PrismML's own llama.cpp fork while the upstream PR sits open. Why it matters: "open weights, runtime-gated" is the honest frame: the weights are free, the runnable path is the vendor's fork, and the 98.2% quality-retention figure is PrismML's claim, not an independent measurement. (Simon Willison)
NEAR AI Cloud's confidential inference picks up non-X primaries
The SayGm route (Bittensor subnet 28) is now documented outside X: NEAR AI's blog describes an Intel TDX gateway enclave routing into TDX VMs with NVIDIA confidential GPUs, and a public attestation endpoint verifiable with a caller-supplied nonce. The vendor's own caveat holds — attestation proves what runs in the enclave at ask-time, not that a given routed request landed there. Why it matters: the distinction is exactly what builders get wrong when they read "verifiable inference," and all privacy-strength framing remains a provider claim. (NEAR AI blog)
NVIDIA publishes its confidential-computing overhead methodology
A NVIDIA technical blog post documents the CC-on versus CC-off measurement methodology and reports DeepSeek-R1 on eight B200s keeping over 96% of output-token throughput at under 5% per-token latency cost with confidential computing on. The figures are NVIDIA-measured on NVIDIA hardware for one workload. Why it matters: it is the best-documented vendor answer yet to "how much does TEE-gated inference cost" — and the reproducible move is running the same comparison on your own workload. (NVIDIA Technical Blog)
Watching tomorrow
OpenAI's legacy-snapshot cutoff lands September 28 — the day after the Sonar retirement — so check both calendars before the weekend ends if your fallback chains cross either provider.
Sources
- Sonar API quickstart — 'Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026.' — Perplexity docs
- The Perplexity Sonar API Retirement: What Changes — LLM Gateway
- OpenRouter: from Seed to Stripe (with Alex Atallah and Anjney Midha) — Latent Space
- Stripe newsroom — Stripe
- Simon Willison — HN comment on Bonsai 2 27B runtime requirements — Simon Willison
- NEAR AI Cloud on SayGm: verifiable confidential inference — NEAR AI
- Enabling Private, High-Performance Production AI Inference with NVIDIA Confidential Computing — NVIDIA Technical Blog
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
