Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3Runtime-tested three-GPU pipeline-parallel deployment of the selective EXL3 Q4 artifact. Each pipeline stage owns complete transformer layers and executes all four independently rotated EXL3 slices locally. Full decode CUDA graphs were captured at batch sizes 1, 2, 4, 8, and 16. A cold-prefix C32 screen delivered 176.82 output tok/s and a warm-prefix C32 screen delivered 203.78 output tok/s. Prefill CUDA graphs remain unsupported for GLM-5.3's hybrid KDA/DSA layers, and the locally built image is not published by digest, so this remains a candidate.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- sglang
- Engine version
- GLM-5.3 selective-EXL3 virtual-slice PP3 overlay
- Graph
- full decode CUDA graphs at batch sizes 1, 2, 4, 8, and 16; full prefill requested but rejected by the runtime because hybrid KDA/DSA layers are not Standard GQA
- Accelerators
- 3
- Tensor parallel
- 1
- Context tokens
- 262,144
- Max concurrency
- 16
- chat
- yes
- reasoning
- yes
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4- Repository
- 0xSero/GLM-5.3-Flash-EXL3-Q4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
sglang
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
glm53-flash-sglang-exl3:pp23-v4
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--model-path/model--served-model-nameglm-5.3-flash-exl3-q4-pp3--tp-size1--pp-size3--ep-size1--quantizationexl3--context-length262144--kv-cache-dtypefp8_e4m3--attention-backenddsa--dsa-prefill-backendflashinfer_sparse_mla--dsa-decode-backendflashinfer_sparse_mla--linear-attn-backendtriton--disable-shared-experts-fusion--chunked-prefill-size256--max-prefill-tokens256--max-running-requests16--max-mamba-cache-size64--pp-max-micro-batch-size8--pp-async-batch-depth1--mem-fraction-static0.88--cuda-graph-config{"decode":{"backend":"full","max_bs":16,"bs":[1,2,4,8,16]},"prefill":{"backend":"full","max_bs":256,"bs":[64,128,256],"full_prefill_max_req":16}}--reasoning-parserglm45--tool-call-parserglm47--host0.0.0.0--port8000
| Flag | Value |
|---|---|
--model-path | /model |
--served-model-name | glm-5.3-flash-exl3-q4-pp3 |
--tp-size | 1 |
--pp-size | 3 |
--ep-size | 1 |
--quantization | exl3 |
--context-length | 262144 |
--kv-cache-dtype | fp8_e4m3 |
--attention-backend | dsa |
--dsa-prefill-backend | flashinfer_sparse_mla |
--dsa-decode-backend | flashinfer_sparse_mla |
--linear-attn-backend | triton |
--chunked-prefill-size | 256 |
--max-prefill-tokens | 256 |
--max-running-requests | 16 |
--max-mamba-cache-size | 64 |
--pp-max-micro-batch-size | 8 |
--pp-async-batch-depth | 1 |
--mem-fraction-static | 0.88 |
--cuda-graph-config | {"decode":{"backend":"full","max_bs":16,"bs":[1,2,4,8,16]},"prefill":{"backend":"full","max_bs":256,"bs":[64,128,256],"full_prefill_max_req":16}} |
--reasoning-parser | glm45 |
--tool-call-parser | glm47 |
--host | 0.0.0.0 |
--port | 8000 |
Environment
| Variable | Value |
|---|---|
SGLANG_EXL3_MAX_BATCH_TOKENS | 256 |
Mounts
| Source | Target |
|---|---|
${MODEL_ROOT}/GLM-5.3-Flash-EXL3-Q4 | /model (read-only) |
${RUNTIME_ASSETS}/chat-template-mm.jinja | /chat-template.jinja (read-only) |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 32 | — | 22.6 | — | accepted-c1-baseline | glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep |
| 8 | 32 | — | 14.7 | — | accepted-shared-prefix | glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep |
| 16 | 32 | — | 7.5 | — | accepted-first-cold-shared-prefix | glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep |
| 32 | 36 | — | 5.5 | — | accepted-cold-unique-prefix-throughput | glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep |
| 32 | 32 | — | 6.4 | — | accepted-warm-shared-prefix-throughput | glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 3
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
- model instance id
- 0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
engine
- graph mode
- full decode CUDA graphs at batch sizes 1, 2, 4, 8, and 16; full prefill requested but rejected by the runtime because hybrid KDA/DSA layers are not Standard GQA
- name
- sglang
- version
- GLM-5.3 selective-EXL3 virtual-slice PP3 overlay
serving
- expert parallel
- 1
- kv cache dtype
- fp8_e4m3
- max concurrency
- 16
- max context tokens
- 262,144
- pipeline parallel
- 3
- tensor parallel
- 1
Provenance & metadata (3)
facts
metadata
- acceptance · all shard load
- Yes
- acceptance · cuda graph capture decode
- Yes
- acceptance · cuda graph capture prefill
- No
- acceptance · endpoint health
- Yes
- acceptance · exl3 kernel initialization
- Yes
- acceptance · generated completion
- Yes
- acceptance · model discovery
- Yes
- benchmark · c16 output tok s
- 120.735
- benchmark · c1 output tok s
- 22.596
- benchmark · c32 cold output tok s
- 176.824
- benchmark · c32 cold total api tok s
- 201.69
- benchmark · c32 warm output tok s
- 203.78
- benchmark · c32 warm total api tok s
- 229.253
- benchmark · c8 output tok s
- 117.369
- checkpoint layout
- glm53-selective-exl3-tp4-v1
- execution layout
- TP1/PP3/EP1; four sealed virtual EXL3 rank slices executed and summed inside each owned pipeline layer
- memory · decode graph free gb by stage after capture
- 24.26, 11.58, 10.4
- memory · prepared weight gb by stage
- 61.67, 73.32, 74.49
- quality note
- The exact runtime returned READY with finish_reason=stop. The artifact's published BF16-to-Q4 KLD remains the quality reference; a new aligned-logit server KLD was not run.
- runtime state
- runtime-tested
- selection
- PP3 accepted for three-GPU capacity and measured throughput. Stock ExLlamaV3 does not currently register GLM-5 Next and cannot consume this custom four-rank-sliced artifact as-is. TP2 remains a separate unaccepted path.
provenance
- captured at
- 2026-08-28T00:58:00Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-28T00:58:00Z | artifact | huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4 ↗ |
| 2026-08-28T00:58:00Z | runtime-acceptance | github.com/0xSero/local-ai-registry ↗ |