GLM-5.3 Open Weights Are Live: Local Deployment, License and Hardware Requirements
Z.ai has released the official GLM-5.3 weights. Its Hugging Face organization now hosts the main zai-org/GLM-5.3 checkpoint plus a separate zai-org/GLM-5.3-BF16 variant, with first-party guidance for local serving.
This changes GLM-5.3 from a hosted-access release into a model that can be evaluated on self-managed infrastructure. The hardware requirement remains substantial: Hugging Face reports roughly 753B parameters for the BF16 repository, placing the official checkpoints in the datacenter/multi-accelerator class.
Quick answer
| Question | Current answer |
|---|---|
| Are the GLM-5.3 weights public? | Yes. Official Z.ai Hugging Face repositories are live |
| Main public checkpoint | zai-org/GLM-5.3, tagged as FP8 on Hugging Face |
| BF16 checkpoint | zai-org/GLM-5.3-BF16 |
| Local serving stacks listed by Z.ai | SGLang, vLLM, TokenSpeed, Transformers, KTransformers and Unsloth |
| Ascend support listed? | Yes; Z.ai points to vLLM-Ascend, xLLM and SGLang paths |
| Hardware class | Multi-accelerator / datacenter by raw checkpoint storage |
| License | Model-specific GLM-5.3 license |
| Thinking control | reasoning_effort supports low, high and max; default is max |
| Best next step | Pin a model revision, verify the license and evaluate on a supported serving stack with sufficient aggregate memory |
What changed with the public checkpoint
At the August 14 launch, Z.ai said the weights would follow after a safety-hardening period. As of August 29, the official Hugging Face artifacts and local-serving instructions are available.
The public release establishes:
- an official downloadable checkpoint;
- a separate BF16 variant;
- FP8 tagging for the main checkpoint;
- documented SGLang, vLLM, TokenSpeed, Transformers, KTransformers and Unsloth serving paths;
- Ascend deployment references;
- a model-specific
glm-5.3license; - documented reasoning-effort and chat-template controls.
Refreshing the original launch guide keeps these deployment details with the page that already owns the GLM-5.3 intent.
Hardware and checkpoint memory
Hugging Face reports about 753B parameters for the BF16 repository. Sparse MoE routing can reduce the compute performed for each token, while checkpoint storage is governed by the full parameter set.
A first-order weight-only estimate is:
| Representation | Approximate weight storage for 753B parameters |
|---|---|
| BF16 / FP16, ~2 bytes per parameter | ~1.51 TB |
| 8-bit / FP8-equivalent, ~1 byte per parameter | ~753 GB |
| Idealized 4-bit packing, ~0.5 byte per parameter | ~377 GB |
These are arithmetic estimates rather than measured end-to-end VRAM requirements. Production serving also allocates memory for KV cache, activations, allocator overhead, communication buffers, runtime workspace, metadata and potentially speculative-decoding state. Quantized formats can add scales, metadata and padding.
The 4-bit row is a storage reference calculation; Z.ai's official release discussed here provides the FP8 and BF16 checkpoint paths.
For capacity planning, GLM-5.3 belongs in multi-accelerator servers or specialized offload configurations rather than the 32B/70B workstation class.
Official local-serving paths
Z.ai's current model card lists:
- SGLang;
- vLLM;
- TokenSpeed;
- Transformers;
- KTransformers;
- Unsloth.
For Huawei Ascend deployments, the card also points to vLLM-Ascend, xLLM and SGLang.
A minimal vLLM path shown through the official integration is conceptually:
vllm serve "zai-org/GLM-5.3"
Production configuration still requires a viable parallelism topology, enough aggregate accelerator memory, a compatible model/runtime revision pair and benchmarking at representative context lengths and concurrency.
FP8 versus BF16
The main zai-org/GLM-5.3 repository is tagged FP8, while Z.ai also publishes zai-org/GLM-5.3-BF16.
- FP8 is the more memory-efficient official starting point when the accelerator and kernels support it well.
- BF16 preserves a higher-precision weight representation and roughly doubles first-order weight storage compared with one-byte-per-parameter FP8 arithmetic.
Kernel support, tensor/expert parallelism, interconnect bandwidth and workload quality should be measured alongside nominal checkpoint size.
Model-specific license
Hugging Face identifies GLM-5.3 with a model-specific glm-5.3 license and marks the repository metadata as license: other with license_name: glm-5.3.
Operators should therefore read the exact checkpoint license before commercial redistribution, hosted-service deployment, modification or embedding the model into another product. This differs from the MIT licensing used by some other Z.ai releases, including GLM-5.3-Flash.
Reasoning controls
GLM-5.3 exposes three reasoning_effort levels:
low;high;max.
Z.ai says max is the default and recommends it for benchmark reproduction. The launch documentation also states that disabling thinking is no longer supported for GLM-5.3 API usage.
The Hugging Face model card adds a chat-template setting: clear_thinking defaults to false, while Z.ai recommends setting clear_thinking=true explicitly for normal chat scenarios.
These settings affect latency, token consumption and conversation behavior, so benchmark and production records should include the exact reasoning configuration.
Coding benchmark evidence
Z.ai's current model card reports:
| Benchmark | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Terminal Bench 2.1 | 88.2 | 81.0 |
| Terminal Bench 3.0 | 28.3 | 4.6 |
| DeepSWE v1.1 | 66.9 | 46.2 |
| SWE-Marathon v1.1 | 42.5 | 19.4 |
| AutomationBench v1.0.6 | 48.2 | 26.2 |
These are Z.ai-reported results. Several evaluations use long timeouts, large context limits, specific agent harnesses or repeated rollouts, so comparative testing should preserve equivalent settings.
A production coding-agent trial should track:
- task completion rate;
- tests passed without manual correction;
- wall-clock time;
- input, cached-input and output tokens;
- retry count;
- tool-call failures;
- generated diff size and defect rate;
- infrastructure cost per accepted task.
Cybersecurity capability and controls
Z.ai reports:
- 84.5% on CyberGym versus 77.2% for GLM-5.2;
- 54.4% on ExploitBench versus 24.4% for GLM-5.2;
- 105 / 130 completed ExploitGym tasks under Z.ai's normalized 2-hour / 6-hour budgets versus 29 / 39 for GLM-5.2.
The model card documents additional evaluation details, including 1,507 CyberGym tasks, 41 ExploitBench tasks across three revisions and network/domain restrictions intended to reduce benchmark leakage.
These remain vendor-reported evaluations. Public weights make it possible for security teams to run the model inside controlled environments, which increases deployment control while also placing a capable coding/security model closer to repositories, shells and security tooling.
Defensive deployments should use explicit target authorization, isolated credentials, scoped filesystem/network/shell permissions, approval gates for destructive or external actions, reproducible finding validation, normal code review and audit logs for high-impact actions.
GLM-5.3 versus GLM-5.3-Flash
GLM-5.3-Flash is a separate model. Z.ai describes Flash as a newly trained 320B-total / 18B-active natively multimodal model with hybrid sparse/linear attention.
GLM-5.3 uses the same base model as GLM-5.2 with additional post-training and a substantially larger public checkpoint by raw parameter storage.
| Model | Deployment profile |
|---|---|
| GLM-5.3 | Larger coding/security-focused capability target; ~753B-class BF16 repository; model-specific license |
| GLM-5.3-Flash | 320B/18B MoE, native multimodality, MIT weights, lower memory class than full GLM-5.3 |
Organizations should choose according to workload quality, multimodal requirements, license and available accelerator memory.
Deployment sequence
- Read the GLM-5.3 license for the intended use.
- Pin the Hugging Face revision rather than floating
mainin production. - Start with FP8 when the accelerator stack supports it and memory is the main constraint.
- Use a first-party-listed serving framework with a compatible release.
- Validate short-context inference before production-scale prompts.
- Measure peak memory as context and concurrency increase.
- Test tool calling and reasoning settings explicitly.
- Run coding/security evaluations inside disposable or isolated environments.
- Record model revision, runtime version, parallelism topology and reasoning configuration with benchmarks.
- Compare infrastructure cost and task quality against smaller alternatives such as GLM-5.3-Flash.
Bottom line
GLM-5.3's official weights and self-hosted serving guidance are now public, making direct deployment evaluation possible. The model remains a datacenter-scale checkpoint, with roughly 753B parameters reported for the BF16 repository and about 1.5 TB of first-order BF16 weight storage before runtime overhead.
The other major operational detail is licensing: GLM-5.3 uses its own model-specific license. Large multi-accelerator environments can now benchmark the model directly; smaller local-AI workstations will generally find GLM-5.3-Flash or other lower-parameter models more practical.
Sources
- Z.ai, GLM-5.3 technical announcement: https://z.ai/blog/glm-5.3
- Z.ai / Hugging Face, GLM-5.3 model card and weights: https://huggingface.co/zai-org/GLM-5.3
- Z.ai / Hugging Face, GLM-5.3 BF16 checkpoint: https://huggingface.co/zai-org/GLM-5.3-BF16
- Z.ai, GLM-5.3-Flash technical announcement: https://z.ai/blog/glm-5.3-flash