Recipe

glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4

glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4

Validated GLM-5.3 Flash NVFP4 vLLM TP4 across four DGX Spark GB10 nodes with DFlash2 speculation, FP8 KV cache, breakable piecewise/full CUDA graphs, a 524,288-token context envelope, uncapped C1 structured matched-window measurements, unified-memory telemetry, and chat, reasoning, tool, structured-output, and vision sidecars.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.1.dev20051+g487ecf187
Graph
breakable CUDA graph with piecewise target capture, full target capture at C1, and full DFlash2 capture at C1
Accelerators
4
Tensor parallel
4
Context tokens
524,288
Max concurrency
1
KV cache tokens
1,816,492
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4
Repository
LibertAIDAI/GLM-5.3-Flash-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

vllm

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/tonyd2wild/vllm-glm53-flash@sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6
Digest
sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6
Compose file
docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-compose.yml
Port
8000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. serve
  2. /model
  3. --served-model-name
  4. glm-5.3-flash
  5. --host
  6. 0.0.0.0
  7. --port
  8. 8000
  9. --trust-remote-code
  10. --tensor-parallel-size
  11. 4
  12. --gpu-memory-utilization
  13. 0.85
  14. --max-model-len
  15. 524288
  16. --max-num-seqs
  17. 1
  18. --max-num-batched-tokens
  19. 2048
  20. --block-size
  21. 2304
  22. --moe-backend
  23. marlin
  24. --speculative-config
  25. {"method":"dflash","model":"/draft","num_speculative_tokens":7}
  26. --kv-cache-dtype
  27. fp8_e4m3
  28. --kv-cache-memory
  29. 12884901888
  30. --tool-call-parser
  31. glm47
  32. --enable-auto-tool-choice
  33. --reasoning-parser
  34. glm45
  35. --chat-template
  36. /chat-template-mm.jinja
  37. --default-chat-template-kwargs
  38. {"enable_thinking":true}
  39. --distributed-executor-backend
  40. mp
  41. --nnodes
  42. 4
  43. --node-rank
  44. ${NODE_RANK}
  45. --master-addr
  46. ${HEAD_FABRIC_ADDRESS}
  47. --master-port
  48. 29521
FlagValue
--served-model-nameglm-5.3-flash
--host0.0.0.0
--port8000
--tensor-parallel-size4
--gpu-memory-utilization0.85
--max-model-len524288
--max-num-seqs1
--max-num-batched-tokens2048
--block-size2304
--moe-backendmarlin
--speculative-config{"method":"dflash","model":"/draft","num_speculative_tokens":7}
--kv-cache-dtypefp8_e4m3
--kv-cache-memory12884901888
--tool-call-parserglm47
--reasoning-parserglm45
--chat-template/chat-template-mm.jinja
--default-chat-template-kwargs{"enable_thinking":true}
--distributed-executor-backendmp
--nnodes4
--node-rank${NODE_RANK}
--master-addr${HEAD_FABRIC_ADDRESS}
--master-port29521

Environment

VariableValue
FLASHINFER_CUDA_ARCH_LIST12.1a
FLASHINFER_DISABLE_VERSION_CHECK1
GLOO_SOCKET_IFNAME${FABRIC_INTERFACE}
HF_HOME/cache/huggingface
HF_HUB_OFFLINE1
MN_IF_NAME${FABRIC_INTERFACE}
NCCL_CROSS_NIC0
NCCL_CUMEM_ENABLE0
NCCL_DEBUGWARN
NCCL_IB_ADDR_FAMILYAF_INET
NCCL_IB_DISABLE0
NCCL_IB_GID_INDEX3
NCCL_IB_HCA${NCCL_IB_HCA}
NCCL_IB_MERGE_NICS0
NCCL_IB_ROCE_VERSION_NUM2
NCCL_IGNORE_CPU_AFFINITY1
NCCL_NETIB
NCCL_NVLS_ENABLE0
NCCL_SOCKET_IFNAME${FABRIC_INTERFACE}
PYTORCH_CUDA_ALLOC_CONFexpandable_segments:True
TORCH_CUDA_ARCH_LIST12.1a
TORCH_NCCL_ASYNC_ERROR_HANDLING1
TP_SOCKET_IFNAME${FABRIC_INTERFACE}
TRANSFORMERS_OFFLINE1
VLLM_ENGINE_READY_TIMEOUT_S3600
VLLM_HOST_IP${NODE_FABRIC_ADDRESS}

Mounts

SourceTarget
${MODEL_ROOT}/GLM-5.3-Flash-NVFP4/model (read-only)
${MODEL_ROOT}/GLM-5.3-Flash-DFlash2/draft (read-only)
asset/glm-5-3-flash-multimodal-chat-template.jinja/chat-template-mm.jinja (read-only)
asset/glm-5-3-flash-vllm-sparse-attn-indexer-kpool.py/usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/sparse_attn_indexer_kpool.py (read-only)
${CACHE_ROOT}/glm53-vllm-512k/cache

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
18,1821,764.9112.14,636.4accepted-sustainedglm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
dgx-spark-gb10-128gb
id
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4
model instance id
libertaidai-glm-5-3-flash-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
breakable CUDA graph with piecewise target capture, full target capture at C1, and full DFlash2 capture at C1
name
vllm
version
0.1.dev20051+g487ecf187

serving

kv cache tokens
1,816,492
max concurrency
1
max context tokens
524,288
tensor parallel
4
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-09-01T06:34:14Z
capabilities.chat · reason live-chat-completion-sidecar-passed
capabilities.chat · state known
capabilities.reasoning · provenance · captured at
2026-09-01T06:34:14Z
capabilities.reasoning · reason live-reasoning-sidecar-returned-separated-reasoning-and-correct-answer
capabilities.reasoning · state known
capabilities.tools · provenance · captured at
2026-09-01T06:34:14Z
capabilities.tools · reason live-forced-tool-call-sidecar-returned-the-required-function-and-arguments
capabilities.tools · state known
capabilities.vision · provenance · captured at
2026-09-01T06:34:14Z
capabilities.vision · reason live-inline-png-sidecar-identified-the-solid-red-image
capabilities.vision · state known
engine.graph mode · provenance · captured at
2026-09-01T06:34:14Z
engine.graph mode · reason startup-log-confirmed-breakable-piecewise-full-target-and-full-dflash2-cuda-graph-capture
engine.graph mode · state known
engine.version · provenance · captured at
2026-09-01T06:34:14Z
engine.version · reason live-version-endpoint-returned-vllm-0.1.dev20051+g487ecf187
engine.version · state known
metadata.runtime identity.published oci manifest digest · provenance · captured at
2026-09-01T06:34:14Z
metadata.runtime identity.published oci manifest digest · reason skopeo-resolved-the-pullable-ghcr-manifest-digest
metadata.runtime identity.published oci manifest digest · state known
serving.kv cache tokens · provenance · captured at
2026-09-01T06:34:14Z
serving.kv cache tokens · reason startup-log-reported-1816492-kv-cache-tokens
serving.kv cache tokens · state known
serving.max concurrency · provenance · captured at
2026-09-01T06:34:14Z
serving.max concurrency · reason scheduler-is-configured-for-max-num-seqs-1-and-all-accepted-measurements-used-c1
serving.max concurrency · state known
serving.max context tokens · provenance · captured at
2026-09-01T06:34:14Z
serving.max context tokens · reason live-model-endpoint-and-startup-log-reported-524288-token-context
serving.max context tokens · state known

metadata

acceptance evidence
docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-20260901.json
content class
structured JSON-schema generation
launch path normalization
Registry launch normalizes live model and chat-template bind targets to plugin-compatible paths without changing arguments or runtime behavior.
measurement · accepted samples
4
measurement · all stream common window
first-token-to-last-token
measurement · cache hit delta
0
measurement · exact streamed token ids
Yes
measurement · median decode tok s
112.118
measurement · median prefill tok s
1,764.873
measurement · median ttft ms
4,636.368
measurement · peak unified memory used bytes
359,068,327,936
measurement · unified memory scope
sum of MemTotal-MemAvailable across four GB10 unified-memory nodes; includes host processes
quality sidecars · arithmetic
Yes
quality sidecars · reasoning
Yes
quality sidecars · structured json
Yes
quality sidecars · tool call
Yes
quality sidecars · vision
Yes
request policy
uncapped natural-stop requests
runtime identity · architecture
Glm5NextForConditionalGeneration
runtime identity · container config id
sha256:35c6f70ffcba62fd67d7b9d4b4e8300ad177201792ce9cdb1ea18fd449bc23b6
runtime identity · draft architecture
DFlash2DraftModel
runtime identity · kv cache dtype
fp8_e4m3
runtime identity · parallelism
four-node-tp4
runtime identity · published oci manifest digest
sha256:4def0ef644cb2e9814136dcffd5e385e21bc594f48f3b292234051904abe85a6
runtime identity · quantization
NVFP4
runtime identity · speculation
dflash2-7
scheduler · block size
2,304
scheduler · kv cache memory bytes per rank
12,884,901,888
scheduler · max num batched tokens
2,048
scheduler · max num seqs
1

provenance

captured at
2026-09-01T06:34:14Z