Recipe

glm53-flash-exl3-q4-rtxpro6000-sglang-tp4

glm53-flash-exl3-q4-rtxpro6000-sglang-tp4

Runtime-tested selective EXL3 Q4 deployment on four RTX PRO 6000 Blackwell GPUs. Routed experts in layers 3-44 use 4.0 bpw EXL3; the backbone remains BF16. The accepted TP4/EP1 path keeps full CUDA graphs and FP8 E4M3 KV. A sustained C256, ISL 131, OSL 1024 screen delivered 102,609.85 output tokens/minute. Tool and vision execution remain unvalidated, and the locally built runtime image is not yet published by digest, so this remains a candidate.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay
Graph
full decode capture planned to 256 concurrent requests (runtime-captured graph buckets through batch 62) and full prefill capture at batch size 1
Accelerators
4
Tensor parallel
4
Context tokens
262,144
Max concurrency
256
chat
yes
reasoning
yes
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
Repository
0xSero/GLM-5.3-Flash-EXL3-Q4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
glm53-flash-sglang-exl3:stage-20260827-r6

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model-path
  2. /model
  3. --served-model-name
  4. glm-5.3-flash-exl3-q4
  5. --tp-size
  6. 4
  7. --ep-size
  8. 1
  9. --quantization
  10. exl3
  11. --context-length
  12. 262144
  13. --kv-cache-dtype
  14. fp8_e4m3
  15. --attention-backend
  16. dsa
  17. --dsa-prefill-backend
  18. flashinfer_sparse_mla
  19. --dsa-decode-backend
  20. flashinfer_sparse_mla
  21. --linear-attn-backend
  22. triton
  23. --disable-shared-experts-fusion
  24. --chunked-prefill-size
  25. 256
  26. --max-prefill-tokens
  27. 256
  28. --max-running-requests
  29. 256
  30. --mem-fraction-static
  31. 0.80
  32. --cuda-graph-backend-decode
  33. full
  34. --cuda-graph-backend-prefill
  35. full
  36. --cuda-graph-max-bs-decode
  37. 256
  38. --cuda-graph-max-bs-prefill
  39. 1
  40. --reasoning-parser
  41. glm45
  42. --tool-call-parser
  43. glm47
  44. --host
  45. 127.0.0.1
  46. --port
  47. 8000
FlagValue
--model-path/model
--served-model-nameglm-5.3-flash-exl3-q4
--tp-size4
--ep-size1
--quantizationexl3
--context-length262144
--kv-cache-dtypefp8_e4m3
--attention-backenddsa
--dsa-prefill-backendflashinfer_sparse_mla
--dsa-decode-backendflashinfer_sparse_mla
--linear-attn-backendtriton
--chunked-prefill-size256
--max-prefill-tokens256
--max-running-requests256
--mem-fraction-static0.80
--cuda-graph-backend-decodefull
--cuda-graph-backend-prefillfull
--cuda-graph-max-bs-decode256
--cuda-graph-max-bs-prefill1
--reasoning-parserglm45
--tool-call-parserglm47
--host127.0.0.1
--port8000

Environment

VariableValue
SGLANG_EXL3_MAX_BATCH_TOKENS256

Mounts

SourceTarget
${MODEL_ROOT}/GLM-5.3-Flash-EXL3-Q4/model (read-only)
${RUNTIME_ASSETS}/chat-template-mm.jinja/chat-template.jinja (read-only)

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
12973.5226accepted-matched-screenglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
12968.5167.1rejected-matched-regressionglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
12966observed-earlier-tp4-ep1-baselineglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
12964.9rejected-speculation-regressionglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
11,024827.71,248accepted-cold-prefillglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
11,022140.3466.8accepted-warm-prefix-prefillglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
1126550.579238.1accepted-cold-c1glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
2128518.563326.6accepted-cold-concurrency-sweepglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
4128295.962.3671.6accepted-cold-concurrency-sweepglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
8128378.355.4699.6accepted-cold-concurrency-sweepglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
16128327.838.21,538.6accepted-cold-concurrency-sweepglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
3212834330.42,299.7accepted-balanced-latency-profileglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
6412836120.93,732.3accepted-short-output-throughput-profileglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
128128371.417.319,334.7rejected-short-output-throughput-regressionglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
64131191.928.92,800accepted-sustained-warm-prefixglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
128131211.628.438,464.4accepted-sustained-warm-prefixglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
256131222.631.668,677.4accepted-max-output-tpm-profileglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
64127295.817.37,432.7rejected-ep4-saturation-regressionglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
6412734520.13,617rejected-single-batch-overlap-regressionglm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
model instance id
0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes

engine

graph mode
full decode capture planned to 256 concurrent requests (runtime-captured graph buckets through batch 62) and full prefill capture at batch size 1
name
sglang
version
0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay

serving

expert parallel
1
kv cache dtype
fp8_e4m3
max concurrency
256
max context tokens
262,144
tensor parallel
4
Provenance & metadata (3)

facts

metadata

acceptance · artifact forward kl bf16 to q4
0.066
acceptance · artifact perplexity delta percent
2.38
acceptance · artifact top1 agreement percent
91.7
acceptance · cuda graph capture
Yes
acceptance · endpoint health
Yes
acceptance · generated completion
Yes
acceptance · model discovery
Yes
benchmark · accepted single stream output tok s
73.619
benchmark · cold 1024 token prefill tok s
827.74
benchmark · matched ep4 candidate output tok s
563.981
benchmark · max sustained output tok s
1,710.164
benchmark · max sustained output tokens per minute
102,609.852
benchmark · max sustained profile
C256, measured ISL 131, requested OSL 1024, warmed shared prefix
benchmark · max sustained total api tokens per minute
115,736.698
benchmark · mtp5 candidate output tok s
64.861
checkpoint layout
glm53-selective-exl3-tp4-v1
controller recipe id
glm-5.3-flash-exl3-q4
quality note
Artifact KLD is held-out BF16-versus-Q4 evaluation. A new full-server aligned-logit KLD run has not been published.
runtime state
runtime-tested
selection
TP4/EP1 accepted; sustained C256 is the max-output-TPM profile; C32/C64 remain the lower-latency profiles; virtual-slice EP4, MTP5, shared-expert fusion, single-batch overlap, and incomplete GLM-5 Next TBO candidates were rejected

provenance

captured at
2026-08-27T22:05:00Z

sources

captured atkindurl
2026-08-27T22:05:00Zartifacthuggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
2026-08-27T22:05:00Zruntime-acceptancegithub.com/0xSero/local-ai-registry