Recipe
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4Validated GLM-5.3 Flash NVFP4 vLLM TP4 across four DGX Spark GB10 nodes with DFlash2 speculation, FP8 KV cache, breakable piecewise/full CUDA graphs, a 524,288-token context envelope, uncapped C1 structured matched-window measurements, unified-memory telemetry, and chat, reasoning, tool, structured-output, and vision sidecars.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- 0.1.dev20051+g487ecf187
- Graph
- breakable CUDA graph with piecewise target capture, full target capture at C1, and full DFlash2 capture at C1
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 524,288
- Max concurrency
- 1
- KV cache tokens
- 1,816,492
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- yes
Hugging Face model card
Identity
https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4- Repository
- LibertAIDAI/GLM-5.3-Flash-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
vllm
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/tonyd2wild/vllm-glm53-flash@sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6- Digest
sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6- Compose file
docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-compose.yml- Port
- 8000
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
serve/model--served-model-nameglm-5.3-flash--host0.0.0.0--port8000--trust-remote-code--tensor-parallel-size4--gpu-memory-utilization0.85--max-model-len524288--max-num-seqs1--max-num-batched-tokens2048--block-size2304--moe-backendmarlin--speculative-config{"method":"dflash","model":"/draft","num_speculative_tokens":7}--kv-cache-dtypefp8_e4m3--kv-cache-memory12884901888--tool-call-parserglm47--enable-auto-tool-choice--reasoning-parserglm45--chat-template/chat-template-mm.jinja--default-chat-template-kwargs{"enable_thinking":true}--distributed-executor-backendmp--nnodes4--node-rank${NODE_RANK}--master-addr${HEAD_FABRIC_ADDRESS}--master-port29521
| Flag | Value |
|---|---|
--served-model-name | glm-5.3-flash |
--host | 0.0.0.0 |
--port | 8000 |
--tensor-parallel-size | 4 |
--gpu-memory-utilization | 0.85 |
--max-model-len | 524288 |
--max-num-seqs | 1 |
--max-num-batched-tokens | 2048 |
--block-size | 2304 |
--moe-backend | marlin |
--speculative-config | {"method":"dflash","model":"/draft","num_speculative_tokens":7} |
--kv-cache-dtype | fp8_e4m3 |
--kv-cache-memory | 12884901888 |
--tool-call-parser | glm47 |
--reasoning-parser | glm45 |
--chat-template | /chat-template-mm.jinja |
--default-chat-template-kwargs | {"enable_thinking":true} |
--distributed-executor-backend | mp |
--nnodes | 4 |
--node-rank | ${NODE_RANK} |
--master-addr | ${HEAD_FABRIC_ADDRESS} |
--master-port | 29521 |
Environment
| Variable | Value |
|---|---|
FLASHINFER_CUDA_ARCH_LIST | 12.1a |
FLASHINFER_DISABLE_VERSION_CHECK | 1 |
GLOO_SOCKET_IFNAME | ${FABRIC_INTERFACE} |
HF_HOME | /cache/huggingface |
HF_HUB_OFFLINE | 1 |
MN_IF_NAME | ${FABRIC_INTERFACE} |
NCCL_CROSS_NIC | 0 |
NCCL_CUMEM_ENABLE | 0 |
NCCL_DEBUG | WARN |
NCCL_IB_ADDR_FAMILY | AF_INET |
NCCL_IB_DISABLE | 0 |
NCCL_IB_GID_INDEX | 3 |
NCCL_IB_HCA | ${NCCL_IB_HCA} |
NCCL_IB_MERGE_NICS | 0 |
NCCL_IB_ROCE_VERSION_NUM | 2 |
NCCL_IGNORE_CPU_AFFINITY | 1 |
NCCL_NET | IB |
NCCL_NVLS_ENABLE | 0 |
NCCL_SOCKET_IFNAME | ${FABRIC_INTERFACE} |
PYTORCH_CUDA_ALLOC_CONF | expandable_segments:True |
TORCH_CUDA_ARCH_LIST | 12.1a |
TORCH_NCCL_ASYNC_ERROR_HANDLING | 1 |
TP_SOCKET_IFNAME | ${FABRIC_INTERFACE} |
TRANSFORMERS_OFFLINE | 1 |
VLLM_ENGINE_READY_TIMEOUT_S | 3600 |
VLLM_HOST_IP | ${NODE_FABRIC_ADDRESS} |
Mounts
| Source | Target |
|---|---|
${MODEL_ROOT}/GLM-5.3-Flash-NVFP4 | /model (read-only) |
${MODEL_ROOT}/GLM-5.3-Flash-DFlash2 | /draft (read-only) |
asset/glm-5-3-flash-multimodal-chat-template.jinja | /chat-template-mm.jinja (read-only) |
asset/glm-5-3-flash-vllm-sparse-attn-indexer-kpool.py | /usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/sparse_attn_indexer_kpool.py (read-only) |
${CACHE_ROOT}/glm53-vllm-512k | /cache |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 8,182 | 1,764.9 | 112.1 | 4,636.4 | accepted-sustained | glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- dgx-spark-gb10-128gb
- id
- glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4
- model instance id
- libertaidai-glm-5-3-flash-nvfp4--nvfp4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- Yes
engine
- graph mode
- breakable CUDA graph with piecewise target capture, full target capture at C1, and full DFlash2 capture at C1
- name
- vllm
- version
- 0.1.dev20051+g487ecf187
serving
- kv cache tokens
- 1,816,492
- max concurrency
- 1
- max context tokens
- 524,288
- tensor parallel
- 4
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-09-01T06:34:14Z
- capabilities.reasoning · provenance · captured at
- 2026-09-01T06:34:14Z
- capabilities.tools · provenance · captured at
- 2026-09-01T06:34:14Z
- capabilities.vision · provenance · captured at
- 2026-09-01T06:34:14Z
- engine.graph mode · provenance · captured at
- 2026-09-01T06:34:14Z
- engine.version · provenance · captured at
- 2026-09-01T06:34:14Z
- metadata.runtime identity.published oci manifest digest · provenance · captured at
- 2026-09-01T06:34:14Z
- serving.kv cache tokens · provenance · captured at
- 2026-09-01T06:34:14Z
- serving.max concurrency · provenance · captured at
- 2026-09-01T06:34:14Z
- serving.max context tokens · provenance · captured at
- 2026-09-01T06:34:14Z
metadata
- acceptance evidence
- docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-20260901.json
- content class
- structured JSON-schema generation
- launch path normalization
- Registry launch normalizes live model and chat-template bind targets to plugin-compatible paths without changing arguments or runtime behavior.
- measurement · accepted samples
- 4
- measurement · all stream common window
- first-token-to-last-token
- measurement · cache hit delta
- 0
- measurement · exact streamed token ids
- Yes
- measurement · median decode tok s
- 112.118
- measurement · median prefill tok s
- 1,764.873
- measurement · median ttft ms
- 4,636.368
- measurement · peak unified memory used bytes
- 359,068,327,936
- measurement · unified memory scope
- sum of MemTotal-MemAvailable across four GB10 unified-memory nodes; includes host processes
- quality sidecars · arithmetic
- Yes
- quality sidecars · reasoning
- Yes
- quality sidecars · structured json
- Yes
- quality sidecars · tool call
- Yes
- quality sidecars · vision
- Yes
- request policy
- uncapped natural-stop requests
- runtime identity · architecture
- Glm5NextForConditionalGeneration
- runtime identity · container config id
- sha256:35c6f70ffcba62fd67d7b9d4b4e8300ad177201792ce9cdb1ea18fd449bc23b6
- runtime identity · draft architecture
- DFlash2DraftModel
- runtime identity · kv cache dtype
- fp8_e4m3
- runtime identity · parallelism
- four-node-tp4
- runtime identity · published oci manifest digest
- sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6
- runtime identity · quantization
- NVFP4
- runtime identity · speculation
- dflash2-7
- scheduler · block size
- 2,304
- scheduler · kv cache memory bytes per rank
- 12,884,901,888
- scheduler · max num batched tokens
- 2,048
- scheduler · max num seqs
- 1
provenance
- captured at
- 2026-09-01T06:34:14Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-09-01T06:34:14Z | live-four-node-runtime-acceptance | github.com/0xSero/local-ai-registry/blob/main/docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-20260901.json ↗ |
| 2026-09-01T06:34:14Z | published-oci-manifest | github.com/tonyd2wild/vllm-glm53-flash/pkgs/container/vllm-glm53-flash ↗ |