Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Self-Hosting & Hardware

Google, OpenAI, Azure, Snowflake retire hosted models in mid-October

Google retires 16 open-model endpoints off Vertex on Oct 21 and OpenAI shuts 12 legacy GPT snapshots Oct 23; Azure and Snowflake cull models the same week.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A single decommissioned server rack with empty drive bays and loosely coiled network cables stands in a plain datacenter hallway, for the mid-October wave in which Google, OpenAI, Microsoft Azure, and Snowflake retire hosted model endpoints within days of each other. Illustration: DeAI
A single decommissioned server rack with empty drive bays and loosely coiled network cables stands in a plain datacenter hallway, for the mid-October wave in which Google, OpenAI, Microsoft Azure, and Snowflake retire hosted model endpoints within days of each other. Illustration: DeAI

Google Cloud is retiring all 16 of its open-model Model-as-a-Service endpoints on Vertex AI on October 21, and OpenAI is shutting down 12 legacy GPT snapshots on October 23, with Microsoft Azure and Snowflake Cortex culling models in the same window. For builders on affected endpoints, the practical decision is where the workload lands, and for the open-weight half of the list the answer Google's own documentation gives is self-deployment.

Key facts

  • Google Cloud's deprecations page lists 16 open-model MaaS endpoints — DeepSeek-OCR, DeepSeek-R1-0528, DeepSeek-V3.1, DeepSeek-V3.2, GLM 5, GLM 4.7, gpt-oss-20b, Kimi K2 Thinking, Llama 3.3 70B, MiniMax M2, two multilingual E5 embedding models, and five Qwen3 variants — all with the same deprecation date (July 21, 2026) and the same retirement date (October 21, 2026).
  • For each of the 16, Google's recommended alternative is to self-deploy on Model Garden; the GLM 5 and GLM 4.7 Garden targets are live.
  • OpenAI's deprecations page lists 12 legacy model snapshots shutting down October 23, 2026 — gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-1106-preview, gpt-4-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1, o1-pro, o3-mini, and o4-mini, plus the ft-o4-mini fine-tune — replaced by GPT-5.6 Sol, Terra, or Luna.
  • The same October 23 table adds five fine-tuned model lines (ft-gpt-3.5-turbo, ft-gpt-4, ft-gpt-4.1-nano-2025-04-14, ft-babbage-002, ft-davinci-002) shutting down with their bases, per OpenAI's deprecations documentation.
  • Microsoft's Foundry retirement schedule shows Moonshot AI's Kimi-K2.7-Code already retired October 3, 2026, and gpt-4.1-nano retiring October 14, 2026.
  • Snowflake's documentation lists claude-4-sonnet and openai-gpt-4.1 at end-of-life October 14, 2026.

What happened

Each of the four shutdowns is documented on the vendor's own lifecycle page, and the dates line up inside two weeks.

Google Cloud's open-model deprecations page, last updated October 1, 2026, puts every one of the 16 MaaS endpoints on an identical clock: deprecated July 21, retired October 21. That is a three-month notice window. The list spans DeepSeek, Z.ai's GLM, OpenAI's own gpt-oss, Moonshot AI's Kimi, Meta's Llama, MiniMax, Qwen, and the multilingual E5 embedding models. The page's stated alternative is the same in every row: self-deploy on Model Garden, or migrate to another managed endpoint before the retirement date. A retired endpoint, per the page's own definitions, is permanently deactivated — API requests calling the retired model ID fail.

OpenAI's deprecations page dates the April 22, 2026 announcement of the legacy-snapshot cull, which lands October 23: gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-1106-preview, gpt-4-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1-2024-12-17, o1-pro-2025-03-19, o3-mini-2025-01-31, and o4-mini-2025-04-16, plus ft-o4-mini-2025-04-16. Five fine-tuned lines go with them. The page's recommended replacements route to GPT-5.6 Sol, Terra, or Luna, or to GPT-Image-2.5 Sunburst and Flare for the image model — six-month notice for the generally available set. OpenAI's notice-period policy commits to at least six months for GA models and at least three months for specialized variants.

Microsoft's Foundry retirement schedule, last updated September 21, 2026, shows Moonshot AI's Kimi-K2.7-Code already retired from Azure Foundry on October 3, 2026, and gpt-4.1-nano retiring October 14 — the same snapshot OpenAI's own API drops nine days later, on a different date and a different page.

Snowflake's release notes show claude-4-sonnet and openai-gpt-4.1 deprecated August 12 with an end-of-life date of October 14, 2026, in AI_COMPLETE, CORTEX.COMPLETE, the Agents API, and Snowflake CoWork.

No vendor in this set has framed the timing as part of a coordinated strategy, and DeAI News makes no pattern claim. What the pages show is four independent calendars converging on the same window — and, for builders on those endpoints, the same two-week scramble.

Why it matters

The deprecation calendar is the real price list for the hosted convenience tier. A workload on a managed per-token endpoint carries a lifecycle clock alongside its token bill, and this window shows what that clock looks like in practice: three months of notice on the Vertex side, six on OpenAI's, and a retirement that leaves either a migration or a self-deployment as the only ways to keep the pipeline running.

For the open-weight half of the Google list, the outcome is asymmetric. The documented migration path for every one of those 16 endpoints is self-deployment on Model Garden — Google retires the convenience, and the model itself is still deployable because the weights are open and the Garden target pages are live. A closed-snapshot cull like OpenAI's October 23 row has no such path: gpt-4-0613 and o3-mini stop existing at the API boundary, and the only forward is the recommended replacement.

That is the recurring shape of this beat. DeAI covered Perplexity's retirement of its chat completions endpoint, where a provider-level API change landed on downstream builders with a fixed clock, and the DeepSeek V4 Flash silent checkpoint swap, where the model behind a stable API name changed without the name moving. Both are lifecycle risk showing up as operational cost. The self-hosting versus API cost breakdown treats depreciation of convenience as part of the API's effective price; this window is that argument landing on a specific set of dates.

The pattern most worth naming is the one visible in the rows themselves. Google's page retires third-party and open models while keeping its first-party Gemini endpoints; OpenAI routes every replacement to its own current family. A vendor culling third-party hosted endpoints to consolidate on its own API is an observable move in the marketplace. Whether it is a durable strategy shift is something no vendor document here states, so it stays an open question rather than a fact.

Background

Hosted endpoints for open-weight models were the compromise that let builders run DeepSeek, Qwen, Kimi, or GLM without owning hardware: the vendor absorbed the GPU fleet, the scaling, and the on-call, and the builder paid per token. The open-weight model explainer covers what those releases actually grant — downloadable weights, permissive or custom licenses, no API dependency. The trade was availability: someone else ran it, so someone else could decide to stop.

The lifecycle math in this window is the compromise unwinding in one direction. A Vertex endpoint for DeepSeek-V3.1 was convenient until July 21, functional until October 21, and after that its continuation requires a cluster — which is the position the self-hosting cost analysis works through in dollars: hardware amortization, utilization, and operations against a per-token bill that vanishes at the vendor's discretion.

Prior DeAI coverage frames the risk from both sides. The Perplexity Sonar retirement showed an API surface disappearing under a production workload. The DeepSeek V4 Flash silent swap showed the quieter version: an endpoint that keeps its name while the model behind it changes. A retirement and a silent swap are different failure modes of the same dependency — you do not control the thing your prompt is pointed at. Open weights are the hedge because the weights survive any vendor decision; the open-weight half of Google's list keeps that hedge, and OpenAI's snapshot cull shows what its absence looks like.

What's next

The dates land in order: Snowflake Cortex EOLs claude-4-sonnet and openai-gpt-4.1 and Azure Foundry retires gpt-4.1-nano on October 14; Google's 16 open-model endpoints go dark October 21; OpenAI's 12 legacy snapshots and five fine-tuned lines shut down October 23. Between now and then, the operational work is a checklist: inventory which model IDs your stack calls, confirm each against the vendor's lifecycle page, and pick per-workload between the recommended replacement, another managed endpoint, or a Model Garden self-deployment where the weights allow it. The third path is the only one of the three that survives the next cull, and its cost is a fleet of your own.

Questions

Which open models is Google retiring from Vertex AI on October 21?
Google Cloud's deprecations page lists 16 open-model Model-as-a-Service endpoints retiring October 21, 2026: DeepSeek-OCR, DeepSeek-R1-0528, DeepSeek-V3.1, DeepSeek-V3.2, GLM 5, GLM 4.7, gpt-oss-20b, Kimi K2 Thinking, Llama 3.3 70B, MiniMax M2, two multilingual E5 embedding models, and five Qwen3 variants.
What OpenAI models shut down on October 23, 2026?
OpenAI's deprecations page lists 12 legacy snapshots shutting down October 23, 2026, including gpt-3.5-turbo-0125, gpt-4-0613, gpt-4-turbo, gpt-4.1-nano, gpt-4o-2024-05-13, gpt-image-1, o1, o1-pro, and o3-mini, plus five fine-tuned variants. Replacements are GPT-5.6 Sol, Terra, and Luna.
What happens to my Vertex AI calls after October 21, 2026?
Google Cloud's documentation says a retired endpoint is permanently deactivated: API requests calling the retired model ID fail. The documented alternatives are migrating to another managed endpoint or self-deploying the model on Model Garden.
Are the weights for the retired open models still available?
For the open-weight models on the Google list, yes. Model Garden keeps self-deploy links for each one, and every model named has published weights on Hugging Face. What ends on October 21 is Google's hosted per-token endpoint, not access to the model itself.
How much notice do you get before a hosted model disappears?
In this wave, Google deprecated the 16 Vertex endpoints on July 21, 2026, giving three months. OpenAI announced the October 23 GPT snapshot shutdowns on April 22, 2026, six months out. Microsoft's policy guarantees six months for GA models and three for specialized variants.

Sources

  1. Open model deprecations | Gemini Enterprise Agent Platform | Google Cloud Documentation — Google Cloud
  2. Deprecations | OpenAI API — OpenAI
  3. Model retirement schedule - Microsoft Foundry | Microsoft Learn — Microsoft
  4. Cortex model deprecations for August 2026 - Snowflake Documentation — Snowflake

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →