Recipe
glm53-flash-nvfp4-rtxpro6000-sglang-tp4
glm53-flash-nvfp4-rtxpro6000-sglang-tp4Live four-GPU GLM-5.3 Flash configuration with completion, tool, vision, video, context, concurrency, and speed evidence. Candidate until the locally tagged runtime image is published by digest.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- sglang
- Engine version
- lmsysorg/sglang:glm-5.3-flash@sha256:3a97bd50034ca60c6e6c86b8e36a73675d261f6a5eb71197796aee5175409290
- Graph
- full decode capture bs<=8 with NEXTN target-verify graphs
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 1,048,576
- Max concurrency
- 8
- KV cache tokens
- 15,156,736
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- yes
Hugging Face model card
Identity
https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4- Repository
- LibertAIDAI/GLM-5.3-Flash-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
sglang
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
glm53-flash-nvfp4-sglang:exact-20260826- Port
- 8000
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--model-path/model--served-model-nameglm-5.3-flash--tp-size4--ep-size4--context-length1048576--quantizationmodelopt_fp4--attention-backenddsa--dsa-prefill-backendflashinfer_sparse_mla--dsa-decode-backendflashinfer_sparse_mla--linear-attn-backendtriton--kv-cache-dtypefp8_e4m3--moe-runner-backendflashinfer_cutlass--disable-shared-experts-fusion--chunked-prefill-size8192--max-prefill-tokens8192--max-running-requests8--mem-fraction-static0.93--cuda-graph-max-bs-decode8--speculative-algorithmNEXTN--speculative-num-steps5--speculative-eagle-topk1--speculative-num-draft-tokens6--speculative-adaptive--reasoning-parserglm45--tool-call-parserglm47--enable-multimodal--chat-template/chat-template.jinja--host0.0.0.0--port30000
| Flag | Value |
|---|---|
--model-path | /model |
--served-model-name | glm-5.3-flash |
--tp-size | 4 |
--ep-size | 4 |
--context-length | 1048576 |
--quantization | modelopt_fp4 |
--attention-backend | dsa |
--dsa-prefill-backend | flashinfer_sparse_mla |
--dsa-decode-backend | flashinfer_sparse_mla |
--linear-attn-backend | triton |
--kv-cache-dtype | fp8_e4m3 |
--moe-runner-backend | flashinfer_cutlass |
--chunked-prefill-size | 8192 |
--max-prefill-tokens | 8192 |
--max-running-requests | 8 |
--mem-fraction-static | 0.93 |
--cuda-graph-max-bs-decode | 8 |
--speculative-algorithm | NEXTN |
--speculative-num-steps | 5 |
--speculative-eagle-topk | 1 |
--speculative-num-draft-tokens | 6 |
--reasoning-parser | glm45 |
--tool-call-parser | glm47 |
--chat-template | /chat-template.jinja |
--host | 0.0.0.0 |
--port | 30000 |
Environment
| Variable | Value |
|---|---|
CUDA_DEVICE_ORDER | PCI_BUS_ID |
FLASHINFER_CUDA_ARCH_LIST | 12.0f |
SGLANG_FP8_PAGED_MQA_LOGITS_TORCH | 1 |
SGLANG_OPT_DEEPGEMM_HC_PRENORM | 0 |
SGLANG_OPT_FP8_WO_A_GEMM | 0 |
SGLANG_OPT_USE_TILELANG_INDEXER | 1 |
SGLANG_OPT_USE_TOPK_V2 | 0 |
TORCH_CUDA_ARCH_LIST | 12.0a |
Mounts
| Source | Target |
|---|---|
${MODEL_ROOT}/GLM-5.3-Flash-NVFP4 | /model (read-only) |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | — | 141.8 | 124 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 2 | — | — | 123 | 252 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 4 | — | — | 100.3 | 256 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 8 | — | — | 72.5 | 332 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 2,125 | 5,344 | — | 398 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 6,325 | 6,383 | — | 991 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 21,407 | 6,275 | — | 3,411 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 53,315 | 6,180 | — | 8,627 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 106,726 | 6,029 | — | 17,702 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
| 1 | 213,094 | 5,783 | — | 36,850 | observed | glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm53-flash-nvfp4-rtxpro6000-sglang-tp4
- model instance id
- libertaidai-glm-5-3-flash-nvfp4--nvfp4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- Yes
engine
- graph mode
- full decode capture bs<=8 with NEXTN target-verify graphs
- name
- sglang
- version
- lmsysorg/sglang:glm-5.3-flash@sha256:3a97bd50034ca60c6e6c86b8e36a73675d261f6a5eb71197796aee5175409290
serving
- kv cache tokens
- 15,156,736
- max concurrency
- 8
- max context tokens
- 1,048,576
- tensor parallel
- 4
Provenance & metadata (3)
facts
metadata
- acceptance · completion
- Yes
- acceptance · tools
- Yes
- acceptance · video
- Yes
- acceptance · vision
- Yes
- runtime pin missing
- launch.image digest
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |