Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4Runtime-tested selective EXL3 Q4 deployment on four RTX PRO 6000 Blackwell GPUs. Routed experts in layers 3-44 use 4.0 bpw EXL3; the backbone remains BF16. The accepted TP4/EP1 path keeps full CUDA graphs and FP8 E4M3 KV. A sustained C256, ISL 131, OSL 1024 screen delivered 102,609.85 output tokens/minute. Tool and vision execution remain unvalidated, and the locally built runtime image is not yet published by digest, so this remains a candidate.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- sglang
- Engine version
- 0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay
- Graph
- full decode capture planned to 256 concurrent requests (runtime-captured graph buckets through batch 62) and full prefill capture at batch size 1
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 262,144
- Max concurrency
- 256
- chat
- yes
- reasoning
- yes
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4- Repository
- 0xSero/GLM-5.3-Flash-EXL3-Q4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
sglang
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
glm53-flash-sglang-exl3:stage-20260827-r6
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--model-path/model--served-model-nameglm-5.3-flash-exl3-q4--tp-size4--ep-size1--quantizationexl3--context-length262144--kv-cache-dtypefp8_e4m3--attention-backenddsa--dsa-prefill-backendflashinfer_sparse_mla--dsa-decode-backendflashinfer_sparse_mla--linear-attn-backendtriton--disable-shared-experts-fusion--chunked-prefill-size256--max-prefill-tokens256--max-running-requests256--mem-fraction-static0.80--cuda-graph-backend-decodefull--cuda-graph-backend-prefillfull--cuda-graph-max-bs-decode256--cuda-graph-max-bs-prefill1--reasoning-parserglm45--tool-call-parserglm47--host127.0.0.1--port8000
| Flag | Value |
|---|---|
--model-path | /model |
--served-model-name | glm-5.3-flash-exl3-q4 |
--tp-size | 4 |
--ep-size | 1 |
--quantization | exl3 |
--context-length | 262144 |
--kv-cache-dtype | fp8_e4m3 |
--attention-backend | dsa |
--dsa-prefill-backend | flashinfer_sparse_mla |
--dsa-decode-backend | flashinfer_sparse_mla |
--linear-attn-backend | triton |
--chunked-prefill-size | 256 |
--max-prefill-tokens | 256 |
--max-running-requests | 256 |
--mem-fraction-static | 0.80 |
--cuda-graph-backend-decode | full |
--cuda-graph-backend-prefill | full |
--cuda-graph-max-bs-decode | 256 |
--cuda-graph-max-bs-prefill | 1 |
--reasoning-parser | glm45 |
--tool-call-parser | glm47 |
--host | 127.0.0.1 |
--port | 8000 |
Environment
| Variable | Value |
|---|---|
SGLANG_EXL3_MAX_BATCH_TOKENS | 256 |
Mounts
| Source | Target |
|---|---|
${MODEL_ROOT}/GLM-5.3-Flash-EXL3-Q4 | /model (read-only) |
${RUNTIME_ASSETS}/chat-template-mm.jinja | /chat-template.jinja (read-only) |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 29 | — | 73.5 | 226 | accepted-matched-screen | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 29 | — | 68.5 | 167.1 | rejected-matched-regression | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 29 | — | 66 | — | observed-earlier-tp4-ep1-baseline | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 29 | — | 64.9 | — | rejected-speculation-regression | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 1,024 | 827.7 | — | 1,248 | accepted-cold-prefill | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 1,022 | 140.3 | — | 466.8 | accepted-warm-prefix-prefill | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 1 | 126 | 550.5 | 79 | 238.1 | accepted-cold-c1 | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 2 | 128 | 518.5 | 63 | 326.6 | accepted-cold-concurrency-sweep | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 4 | 128 | 295.9 | 62.3 | 671.6 | accepted-cold-concurrency-sweep | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 8 | 128 | 378.3 | 55.4 | 699.6 | accepted-cold-concurrency-sweep | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 16 | 128 | 327.8 | 38.2 | 1,538.6 | accepted-cold-concurrency-sweep | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 32 | 128 | 343 | 30.4 | 2,299.7 | accepted-balanced-latency-profile | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 64 | 128 | 361 | 20.9 | 3,732.3 | accepted-short-output-throughput-profile | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 128 | 128 | 371.4 | 17.3 | 19,334.7 | rejected-short-output-throughput-regression | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 64 | 131 | 191.9 | 28.9 | 2,800 | accepted-sustained-warm-prefix | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 128 | 131 | 211.6 | 28.4 | 38,464.4 | accepted-sustained-warm-prefix | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 256 | 131 | 222.6 | 31.6 | 68,677.4 | accepted-max-output-tpm-profile | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 64 | 127 | 295.8 | 17.3 | 7,432.7 | rejected-ep4-saturation-regression | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
| 64 | 127 | 345 | 20.1 | 3,617 | rejected-single-batch-overlap-regression | glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
- model instance id
- 0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
engine
- graph mode
- full decode capture planned to 256 concurrent requests (runtime-captured graph buckets through batch 62) and full prefill capture at batch size 1
- name
- sglang
- version
- 0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay
serving
- expert parallel
- 1
- kv cache dtype
- fp8_e4m3
- max concurrency
- 256
- max context tokens
- 262,144
- tensor parallel
- 4
Provenance & metadata (3)
facts
metadata
- acceptance · artifact forward kl bf16 to q4
- 0.066
- acceptance · artifact perplexity delta percent
- 2.38
- acceptance · artifact top1 agreement percent
- 91.7
- acceptance · cuda graph capture
- Yes
- acceptance · endpoint health
- Yes
- acceptance · generated completion
- Yes
- acceptance · model discovery
- Yes
- benchmark · accepted single stream output tok s
- 73.619
- benchmark · cold 1024 token prefill tok s
- 827.74
- benchmark · matched ep4 candidate output tok s
- 563.981
- benchmark · max sustained output tok s
- 1,710.164
- benchmark · max sustained output tokens per minute
- 102,609.852
- benchmark · max sustained profile
- C256, measured ISL 131, requested OSL 1024, warmed shared prefix
- benchmark · max sustained total api tokens per minute
- 115,736.698
- benchmark · mtp5 candidate output tok s
- 64.861
- checkpoint layout
- glm53-selective-exl3-tp4-v1
- controller recipe id
- glm-5.3-flash-exl3-q4
- quality note
- Artifact KLD is held-out BF16-versus-Q4 evaluation. A new full-server aligned-logit KLD run has not been published.
- runtime state
- runtime-tested
- selection
- TP4/EP1 accepted; sustained C256 is the max-output-TPM profile; C32/C64 remain the lower-latency profiles; virtual-slice EP4, MTP5, shared-expert fusion, single-batch overlap, and incomplete GLM-5 Next TBO candidates were rejected
provenance
- captured at
- 2026-08-27T22:05:00Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T22:05:00Z | artifact | huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4 ↗ |
| 2026-08-27T22:05:00Z | runtime-acceptance | github.com/0xSero/local-ai-registry ↗ |