Independent/Reader-funded/Infrastructure, not tokens
DeAINEWS

AI you control — open models, private inference, and the networks that run them.

Open-Weights Releases

Run GLM-5.3-Flash Locally with llama.cpp (2026): VRAM, Quants, Setup

Z.ai's 320B GLM-5.3-Flash now runs in llama.cpp as of the Sept 30 merge. GGUF sizes from 92 GB, quant quality data, setup commands, and known limits.

DeAI is powered by Morpheus (mor.org). We cover competing providers on the same terms — see our methodology.

A dual-GPU workstation mid-build on a wooden desk with both cards seated and power cables routed, the class of machine that runs Z.ai's 320B GLM-5.3-Flash from quantized GGUF files in llama.cpp. Illustration: DeAI
A dual-GPU workstation mid-build on a wooden desk with both cards seated and power cables routed, the class of machine that runs Z.ai's 320B GLM-5.3-Flash from quantized GGUF files in llama.cpp. Illustration: DeAI

GLM-5.3-Flash, Z.ai's 320B-parameter open-weight model with just 18B active per token, merged into llama.cpp mainline on September 30, 2026 — meaning the month-old frontier-class model now runs on your own hardware from a single build, with community GGUF quantizations starting near 92 GB.

Key facts

What actually merged

The llama.cpp pull request, opened by timkhronos on August 26 and merged September 30, adds the glm5-next architecture: 45 layers where 34 use Kimi-K2-style linear attention (KDA, reusing that implementation) and 11 use DeepSeek-style sparse attention (DSA) with a pooled indexer, plus DeepSeek-V4-style mixture-of-experts feed-forward blocks and Manifold-Constrained Hyper-Connections (mHC). Most of those pieces already existed in llama.cpp for other models; the new work was the DSA indexer's pooled scoring, the vision tower (a new glm5v projector, distinct from glm4v), and a hybrid memory layout that keeps the recurrent state and the sparse-attention cache side by side.

Three implementation details matter to people running it:

  1. Roughly 1 GB of small tensors stays unquantized — the DSA indexer, mHC mixers, KDA gates, and MLA low-rank paths. This is why aggressive quants still hold up: the precision-sensitive plumbing is protected by the PR's quantization rules rather than left to the quant recipe.
  2. Multi-sequence serving needs --kv-unified. The hybrid memory path currently requires unified KV cache when you run more than one sequence in parallel. Single-session use is unaffected.
  3. MTP (multi-token prediction) is a separate pending PR (#27917) — the NextN tensors are preserved during conversion, but speculative decoding is not in the merged mainline yet. Speed numbers quoted above do not include MTP.

Which GGUF, and what does the quality data say

Two community repos carried the model through the review window. avar6's GLM-5.3-Flash-BF16-gguf (2,900+ downloads last month) published the first working quants with vision support, converted from the BF16 safetensors. Aes Sedai's GLM-5.3-Flash-GGUF took a mixture-of-experts approach: since the FFN tensors dominate this model's size, they are quantized harder while attention-critical tensors stay at Q8_0. Aes Sedai published perplexity and KL-divergence measurements against the BF16 baseline, which is the data that turns quant choice from guesswork:

QuantSizeQuality vs BF16Notes
Q5_K_M224 GB+0.55% perplexityNear-lossless; 24 GB VRAM + 256 GB RAM territory
Q4_K_M188 GB+1.83% perplexityHigh-quality floor for serious use
IQ4_XS148 GB+6.98% perplexityBest size/quality balance if it fits
IQ3_S116 GB+22.88% perplexityThe common consumer split-offload pick
IQ2_S106 GB+33.37% perplexityExists to prove it fits; visibly degraded

The model card's own guidance fills in the rest: Z.ai ships the checkpoint in BF16 with FP8 and F32 tensors (321B parameters on disk), and the model card notes that thinking budget is controlled by a reasoning_effort parameter (low/high/max, defaulting to max) — which in llama.cpp terms means long thinking traces on hard prompts, so budget generation-token headroom accordingly.

Setup: from zero to first token

The merged mainline means no forks and no special flags beyond the ones below. Build llama.cpp from source at commit 649dcb1 or later (any release after September 30, 2026), then:

1. Pull a quant from Hugging Face. For a dual-GPU or GPU-plus-RAM split, IQ3_S or IQ4_XS is the realistic starting band:

huggingface-cli download AesSedai/GLM-5.3-Flash-GGUF \
  --include "*IQ4_XS*" "*mmproj*" --local-dir ./glm53flash

Grab the mmproj file too if you want vision — it is a separate file in both repos.

2. Run with split mode across GPU and system RAM. The documented consumer configuration is a 24 GB GPU plus 128 GB of DDR5:

llama-server -m glm53flash/GLM-5.3-Flash-IQ4_XS.gguf \
  --mmproj glm53flash/mmproj.glmv5v-f16.gguf \
  -ngl 99 --n-cpu-moe 34 -c 32768 --kv-unified --port 8080

--n-cpu-moe keeps the large expert FFN tensors in system RAM while attention layers live on the GPU — the standard pattern for oversized MoE models, and the reason a 320B model is practical at all on one card. --kv-unified is required whenever you plan to run multiple sequences. Tune context (-c) to your RAM budget: the hybrid attention keeps long-context memory growth moderate relative to full-attention models of this size, but the community-reported checkpoint growth at 90K+ tokens is worth watching if you push toward 256K.

3. Call it like any OpenAI-compatible endpoint. llama-server exposes http://localhost:8080/v1, so the same client code you point at any other host works unchanged — base URL, key (any string for local), and the model slug. Vision requests go through the standard multimodal message format once the mmproj is loaded.

Honest limits

  • It is early. The merge is one day old at writing. Expect rough edges — the Metal fusion baseline needed a follow-up PR the same day, and the maintainers' review thread shows the memory-layout code was reworked twice in the final week. Pin a known-good build rather than tracking master blindly.
  • The consumer-hardware numbers are one anecdote. The 300 tok/s prefill / 6-9 tok/s decode figure is a single tester's report on specific hardware, not a benchmark suite, and it degraded as context filled — a known behavior of the pooled-indexer design that may improve as follow-up optimizations land. Treat all community speed figures as starting points for your own measurement.
  • 1M context is an API-side feature. The hybrid architecture is designed for it, but a local rig's practical context is set by your memory budget, not the model's ceiling.
  • Not your only local option. The model card lists SGLang, vLLM, KTransformers, TokenSpeed, and Unsloth as supported serving paths — llama.cpp is the lowest-dependency route, not the highest-throughput one. For multi-user production serving, vLLM or SGLang on datacenter hardware remains the designed-for path.
  • Deployment decisions have a security dimension now. Anthropic's Frontier Red Team reported on September 29 that GLM-5.3's safeguard layer is bypassable and NIST's CAISI assessment reached a matching finding — so if you expose a self-hosted GLM-5.3 endpoint beyond localhost, treat the exposure surface the way our hardening coverage describes for Ollama: bind to localhost, authenticate, and log.

Local versus API: the same model, two trade-off tables

The self-hosted case for GLM-5.3-Flash is stronger than for most 300B-class models because of the active-parameter count: 18B active means inference-time compute is modest even while total knowledge capacity is frontier-class. You are buying: full control of prompts and logs, no per-token billing, and a model that keeps working when a provider changes terms. You are paying: roughly 96-128 GB of memory for the sensible quants, single-digit tokens/s generation on consumer split-offload setups, day-one tooling maturity, and your own ops burden — the same trade our self-hosting versus API cost math covers generically, now at a new point on the capability curve.

If you self-host for privacy rather than cost, note the distinction our TEE explainer draws: running weights on your own hardware is the strongest data-control posture available; hosted "private" endpoints are policy statements unless backed by attestation.

The model itself remains MIT-licensed — commercial use, modification, and redistribution all permitted — so nothing about local deployment has licensing friction. Check the license field on the exact GGUF repo you pull from; quantized repos inherit the base model's licensing in practice, but the card is the authority.

Questions

Can GLM-5.3-Flash run in llama.cpp?
Yes. Support for the glm5-next architecture merged into llama.cpp mainline on September 30, 2026 (PR #27773), so any build from that commit onward runs the model from GGUF files. Community quantizations are on Hugging Face, starting near 92 GB at IQ2_XXS.
How much VRAM does GLM-5.3-Flash need?
The smallest published GGUF (IQ2_XXS, 2.32 bits per weight) is about 92 GB, and the commonly recommended quality tier IQ3_S is about 116 GB, so plan for roughly 96-128 GB of combined VRAM-plus-RAM. A 24 GB GPU plus 128 GB system RAM is a documented working split-offload configuration.
What is GLM-5.3-Flash?
Z.ai's first natively multimodal GLM-5 series model: 320B total parameters with 18B active per token, combining sparse (DSA) and linear (KDA) attention to cut long-context serving cost, with a 1M-token context window on the official API. Weights are on Hugging Face under the MIT license.
Which GLM-5.3-Flash quantization should I pick?
Aes Sedai's perplexity data shows IQ4_XS at 148 GB losing about 7% quality versus the BF16 baseline, IQ3_S at 116 GB about 23%, and IQ2_S at 106 GB about 33%. If your machine fits IQ4_XS, use it; below that, expect visible degradation on hard reasoning tasks.
Does GLM-5.3-Flash vision work in llama.cpp?
Yes, through the mtmd multimodal path with a glm5v vision projector (mmproj) file. Early mmproj builds had issues that were fixed during the PR review; use a recent build and the updated mmproj from the GGUF repos. Multimodal serving is also available through vLLM and SGLang.

Sources

  1. add GLM-5.3-Flash (GLM5-Next) support · llama.cpp PR #27773 — ggml-org on GitHub
  2. zai-org/GLM-5.3-Flash model card — Hugging Face
  3. AesSedai/GLM-5.3-Flash-GGUF (perplexity and KLD table) — Hugging Face
  4. avar6/GLM-5.3-Flash-BF16-gguf — Hugging Face
  5. GLM-5.3-Flash/FlashX overview — Z.ai developer docs — Z.ai
  6. GLM-5: from Vibe Coding to Agentic Engineering (technical report) — arXiv

About DeAI

DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.

Powered by Morpheus and StrandCMS

Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more about the Morpheus Inference API →

Sponsor disclosure — not editorial

Powered by Morpheus and StrandCMS. Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider. StrandCMS is the open-source, agent-first framework this site is built on.

Learn more →