Recipe

glm53-flash-exl3-q4-rtxpro6000-sglang-pp3

glm53-flash-exl3-q4-rtxpro6000-sglang-pp3

Runtime-tested three-GPU pipeline-parallel deployment of the selective EXL3 Q4 artifact. Each pipeline stage owns complete transformer layers and executes all four independently rotated EXL3 slices locally. Full decode CUDA graphs were captured at batch sizes 1, 2, 4, 8, and 16. A cold-prefix C32 screen delivered 176.82 output tok/s and a warm-prefix C32 screen delivered 203.78 output tok/s. Prefill CUDA graphs remain unsupported for GLM-5.3's hybrid KDA/DSA layers, and the locally built image is not published by digest, so this remains a candidate.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
GLM-5.3 selective-EXL3 virtual-slice PP3 overlay
Graph
full decode CUDA graphs at batch sizes 1, 2, 4, 8, and 16; full prefill requested but rejected by the runtime because hybrid KDA/DSA layers are not Standard GQA
Accelerators
3
Tensor parallel
1
Context tokens
262,144
Max concurrency
16
chat
yes
reasoning
yes
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
Repository
0xSero/GLM-5.3-Flash-EXL3-Q4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
glm53-flash-sglang-exl3:pp23-v4

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model-path
  2. /model
  3. --served-model-name
  4. glm-5.3-flash-exl3-q4-pp3
  5. --tp-size
  6. 1
  7. --pp-size
  8. 3
  9. --ep-size
  10. 1
  11. --quantization
  12. exl3
  13. --context-length
  14. 262144
  15. --kv-cache-dtype
  16. fp8_e4m3
  17. --attention-backend
  18. dsa
  19. --dsa-prefill-backend
  20. flashinfer_sparse_mla
  21. --dsa-decode-backend
  22. flashinfer_sparse_mla
  23. --linear-attn-backend
  24. triton
  25. --disable-shared-experts-fusion
  26. --chunked-prefill-size
  27. 256
  28. --max-prefill-tokens
  29. 256
  30. --max-running-requests
  31. 16
  32. --max-mamba-cache-size
  33. 64
  34. --pp-max-micro-batch-size
  35. 8
  36. --pp-async-batch-depth
  37. 1
  38. --mem-fraction-static
  39. 0.88
  40. --cuda-graph-config
  41. {"decode":{"backend":"full","max_bs":16,"bs":[1,2,4,8,16]},"prefill":{"backend":"full","max_bs":256,"bs":[64,128,256],"full_prefill_max_req":16}}
  42. --reasoning-parser
  43. glm45
  44. --tool-call-parser
  45. glm47
  46. --host
  47. 0.0.0.0
  48. --port
  49. 8000
FlagValue
--model-path/model
--served-model-nameglm-5.3-flash-exl3-q4-pp3
--tp-size1
--pp-size3
--ep-size1
--quantizationexl3
--context-length262144
--kv-cache-dtypefp8_e4m3
--attention-backenddsa
--dsa-prefill-backendflashinfer_sparse_mla
--dsa-decode-backendflashinfer_sparse_mla
--linear-attn-backendtriton
--chunked-prefill-size256
--max-prefill-tokens256
--max-running-requests16
--max-mamba-cache-size64
--pp-max-micro-batch-size8
--pp-async-batch-depth1
--mem-fraction-static0.88
--cuda-graph-config{"decode":{"backend":"full","max_bs":16,"bs":[1,2,4,8,16]},"prefill":{"backend":"full","max_bs":256,"bs":[64,128,256],"full_prefill_max_req":16}}
--reasoning-parserglm45
--tool-call-parserglm47
--host0.0.0.0
--port8000

Environment

VariableValue
SGLANG_EXL3_MAX_BATCH_TOKENS256

Mounts

SourceTarget
${MODEL_ROOT}/GLM-5.3-Flash-EXL3-Q4/model (read-only)
${RUNTIME_ASSETS}/chat-template-mm.jinja/chat-template.jinja (read-only)

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
13222.6accepted-c1-baselineglm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
83214.7accepted-shared-prefixglm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
16327.5accepted-first-cold-shared-prefixglm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
32365.5accepted-cold-unique-prefix-throughputglm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
32326.4accepted-warm-shared-prefix-throughputglm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
3
hardware id
rtx-pro-6000-blackwell-96gb
id
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
model instance id
0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes

engine

graph mode
full decode CUDA graphs at batch sizes 1, 2, 4, 8, and 16; full prefill requested but rejected by the runtime because hybrid KDA/DSA layers are not Standard GQA
name
sglang
version
GLM-5.3 selective-EXL3 virtual-slice PP3 overlay

serving

expert parallel
1
kv cache dtype
fp8_e4m3
max concurrency
16
max context tokens
262,144
pipeline parallel
3
tensor parallel
1
Provenance & metadata (3)

facts

metadata

acceptance · all shard load
Yes
acceptance · cuda graph capture decode
Yes
acceptance · cuda graph capture prefill
No
acceptance · endpoint health
Yes
acceptance · exl3 kernel initialization
Yes
acceptance · generated completion
Yes
acceptance · model discovery
Yes
benchmark · c16 output tok s
120.735
benchmark · c1 output tok s
22.596
benchmark · c32 cold output tok s
176.824
benchmark · c32 cold total api tok s
201.69
benchmark · c32 warm output tok s
203.78
benchmark · c32 warm total api tok s
229.253
benchmark · c8 output tok s
117.369
checkpoint layout
glm53-selective-exl3-tp4-v1
execution layout
TP1/PP3/EP1; four sealed virtual EXL3 rank slices executed and summed inside each owned pipeline layer
memory · decode graph free gb by stage after capture
24.26, 11.58, 10.4
memory · prepared weight gb by stage
61.67, 73.32, 74.49
quality note
The exact runtime returned READY with finish_reason=stop. The artifact's published BF16-to-Q4 KLD remains the quality reference; a new aligned-logit server KLD was not run.
runtime state
runtime-tested
selection
PP3 accepted for three-GPU capacity and measured throughput. Stock ExLlamaV3 does not currently register GLM-5 Next and cannot consume this custom four-rank-sliced artifact as-is. TP2 remains a separate unaccepted path.

provenance

captured at
2026-08-28T00:58:00Z

sources

captured atkindurl
2026-08-28T00:58:00Zartifacthuggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
2026-08-28T00:58:00Zruntime-acceptancegithub.com/0xSero/local-ai-registry