Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

BreakingProvider Policy & Trust

Perplexity retires the Sonar chat-completions API

Perplexity's Sonar chat-completions surface went dark: sonar-pro and sonar-reasoning-pro stop being routable Sept 27 with no drop-in successor.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single brushed-metal equipment cabinet with dark unlit screens and a column of solid amber indicator lights, standing alone in a dim datacenter corridor — the serving hardware behind a model id going dark on September 27. Illustration: DeAI
A single brushed-metal equipment cabinet with dark unlit screens and a column of solid amber indicator lights, standing alone in a dim datacenter corridor — the serving hardware behind a model id going dark on September 27. Illustration: DeAI

Perplexity has retired the Sonar chat-completions surface. The perplexity/sonar model id survives: its upstream transport moved to the company's Agent API on September 25, keeping the request shape and the top-level search_results array. But sonar-pro and sonar-reasoning-pro stop being routable on September 27, with no drop-in successor — and that deadline is tomorrow.

Key facts

  • Perplexity's Sonar quickstart banner now reads: "Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026" (Perplexity docs).
  • Two of the three Sonar-era model ids, sonar-pro and sonar-reasoning-pro, are not accepted by the Agent API and return errors from September 27, per gateway operator LLM Gateway.
  • The canonical Agent API endpoint is POST /v1/agent, which takes an input string and a preset — fast, low, medium, high — and returns a typed output array with message and search_results items, per Perplexity's migration guide.
  • Perplexity's own mapping sends Sonar to the fast preset, Sonar Pro to low, Sonar Reasoning Pro to medium, and Sonar Deep Research to high (migration guide).
  • The Agent API web_search tool costs $1.00 per 1,000 invocations on the fast preset, down from $2.50, per the Perplexity changelog.
  • On the old chat-completions endpoint, every request searched by default and the usage object reported a flat request_cost on top of token costs, per the API reference.

What happened

The change arrived in three dated steps. Perplexity announced the migration on August 13, 2026 with a 45-day notice window. On September 25, the transport behind perplexity/sonar switched: the model id stayed, but requests now execute on the Agent API. On September 27, the day after this article's publication date, perplexity/sonar-pro and perplexity/sonar-reasoning-pro stop being routable entirely. Perplexity's own quickstart carries the deadline as a banner rather than a changelog post.

The behavioral shift is bigger than an endpoint rename. Perplexity's migration guide describes the difference plainly: search becomes "a tool the model can call rather than something baked in." On the Sonar chat-completions endpoint, every request grounded against the web and returned a top-level search_results array. On the Agent API, the model decides whether to search. Callers who need the old always-grounded behavior now have to either pick a preset that leans on search or attach and require a web_search tool themselves.

Perplexity's mapping table points each old tier at a preset: fast for Sonar, low for Sonar Pro, medium for Sonar Reasoning Pro. Those presets are not the old models. Per the changelog, the fast preset runs openai/gpt-6-luna with reasoning effort set to none — so a Sonar Pro caller who follows the mapping is served by a different underlying model family. Perplexity's docs claim quality improvements across the mappings ("more accurate than Sonar," "far stronger on step-by-step reasoning"), and the high/Deep Research mapping is described as "often at a lower per-request cost." Those are Perplexity's own benchmark claims, not independently verified results.

On pricing, the mechanics are checkable even where the conclusions are not. The old Sonar endpoint's API reference shows a request_cost field — a flat per-request charge alongside token costs. The Agent API meters search per invocation, at $1.00 per 1,000 web_search calls on the fast preset per the changelog. LLM Gateway, which fronts all three Perplexity model ids, reports that output tokens are priced higher than before, input tokens considerably lower, and cached input is now billed separately — and that a typical grounded question measured about a third cheaper in its own testing. That is one gateway's measurement, not an independent benchmark.

The third-party summaries agree on the dates. DNotifier's write-up documents the same retirement window and the same preset mappings, without adding anything Perplexity's docs lack.

Why it matters

For builders deciding where to run models, this is the useful contrast to the silent swap. We covered DeepSeek's silent retirement of V4-Flash earlier this month: an id that kept working while the checkpoint behind it changed, disclosed only in a pricing-page footnote. Perplexity chose the opposite failure mode. The pro tiers do not quietly serve something else; they error. The caller sees the break, on a date published 45 days in advance.

That visible break has a cost, and it lands on anyone pinned to the retiring ids. A pipeline calling sonar-pro cannot re-point at a new model name and move on — the migration is a rewrite against tool-calling semantics, a different request shape (input plus preset instead of a messages array), and a typed output array instead of chat choices. Evaluations have to be re-run from scratch, because the preset runs a different model with different search behavior. Gateway operators have had to make the same call LLM Gateway describes: attach a required web_search tool to preserve Sonar's always-grounded behavior under the old id, or surface an error rather than silently substitute a preset. Those are the operational questions worth asking of any gateway you use: does it pin the model id you asked for, does it substitute a preset, and does a price change also change what actually runs.

The middle layer is where these decisions get made, and it is under consolidation — Stripe's acquisition of OpenRouter put a payments company in the routing path, which we touched on in our gateway logging coverage and our marketplace comparison. A deprecation that lands at the transport layer, one step below the model ids most dashboards show, is exactly the kind of change that propagates through that layer quietly.

Background

Sonar has been Perplexity's developer-facing brand since 2024, when the company split its hosted Llama-3.1-based online models into the Sonar line (changelog). The chat-completions shape carried through every subsequent generation — Sonar, Sonar Pro, Sonar Reasoning Pro, Sonar Deep Research — with search grounded into every request and citations returned as URL lists plus a structured search_results array.

The retirement is also the last step of a scheduled arc. Perplexity's changelog shows the MCP server moved its model-backed tools from Sonar models to Agent API presets in July 2026, and the Agent API reached general availability in February. By September, the chat-completions surface was the last Sonar transport standing.

The deprecation calendar is crowded right now. OpenAI's legacy-snapshot cutoff on its own API lands September 28 — the day after this retirement executes — and Perplexity's changelog separately schedules the retirement of older OpenAI model ids (openai/gpt-5.4 through openai/gpt-5-mini) on its Agent and Router APIs for October 24. Where your fallback chains point, and whether a gateway passes deprecation notices through, is the difference between a scheduled rewrite and an outage. Our price index and cheap-API roundup track the provider field these deprecations keep reshaping.

What's next

September 27 is the hard date: after it, sonar-pro and sonar-reasoning-pro requests return errors with no drop-in successor, and the practical options are migrating to a preset or moving the traffic to another grounded provider. Check your usage dashboards for the two retiring ids before then. After the cutover, two things are worth watching: whether independent benchmarks of the Agent API presets appear to test Perplexity's quality and cost claims, and whether more gateways follow LLM Gateway's pattern of surfacing a visible error instead of substituting a preset. The next dated item on the calendar is OpenAI's legacy-snapshot cutoff on September 28.

Questions

What is being retired on September 27, 2026?
Perplexity's Sonar chat-completions service ends September 27, 2026. The perplexity/sonar-pro and perplexity/sonar-reasoning-pro model ids stop being routable that day, per Perplexity's docs and gateway documentation. Perplexity's migration guide maps the old tiers to Agent API presets, which run different underlying models.
Does perplexity/sonar still work after the retirement?
The perplexity/sonar model id keeps working: its upstream transport moved to the Agent API on September 25, 2026. The request shape and the top-level search_results array stay the same, but search becomes a tool the model decides to call rather than something every request does.
What replaces Sonar's chat-completions API?
The Agent API, a POST /v1/agent endpoint that takes an input string and a preset (fast, low, medium, high) and returns a typed output array with message and search_results items, per Perplexity's migration guide.
Is the Agent API cheaper than Sonar was?
That is a claim, not a settled fact. Perplexity's docs say the Agent API is more performant and cost-effective; gateway operator LLM Gateway reports a grounded question measuring about a third cheaper in its own testing. No independent benchmark of Agent API presets was available at publication.

Sources

  1. Sonar API quickstart — 'Sonar Chat Completions is now Agent API. Sonar will be supported until September 27, 2026.' — Perplexity docs
  2. Migrate from Sonar to the Agent API — Perplexity docs
  3. Perplexity changelog — Agent API, preset and deprecation entries — Perplexity docs
  4. The Perplexity Sonar API Retirement: What Changes — LLM Gateway
  5. Perplexity's Sonar API Is Retiring | September 2026 — DNotifier

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →