Recipe

glm53-flash-nvfp4-rtxpro6000-sglang-tp4

glm53-flash-nvfp4-rtxpro6000-sglang-tp4

Live four-GPU GLM-5.3 Flash configuration with completion, tool, vision, video, context, concurrency, and speed evidence. Candidate until the locally tagged runtime image is published by digest.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
lmsysorg/sglang:glm-5.3-flash@sha256:3a97bd50034ca60c6e6c86b8e36a73675d261f6a5eb71197796aee5175409290
Graph
full decode capture bs<=8 with NEXTN target-verify graphs
Accelerators
4
Tensor parallel
4
Context tokens
1,048,576
Max concurrency
8
KV cache tokens
15,156,736
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/LibertAIDAI/GLM-5.3-Flash-NVFP4
Repository
LibertAIDAI/GLM-5.3-Flash-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
glm53-flash-nvfp4-sglang:exact-20260826
Port
8000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model-path
  2. /model
  3. --served-model-name
  4. glm-5.3-flash
  5. --tp-size
  6. 4
  7. --ep-size
  8. 4
  9. --context-length
  10. 1048576
  11. --quantization
  12. modelopt_fp4
  13. --attention-backend
  14. dsa
  15. --dsa-prefill-backend
  16. flashinfer_sparse_mla
  17. --dsa-decode-backend
  18. flashinfer_sparse_mla
  19. --linear-attn-backend
  20. triton
  21. --kv-cache-dtype
  22. fp8_e4m3
  23. --moe-runner-backend
  24. flashinfer_cutlass
  25. --disable-shared-experts-fusion
  26. --chunked-prefill-size
  27. 8192
  28. --max-prefill-tokens
  29. 8192
  30. --max-running-requests
  31. 8
  32. --mem-fraction-static
  33. 0.93
  34. --cuda-graph-max-bs-decode
  35. 8
  36. --speculative-algorithm
  37. NEXTN
  38. --speculative-num-steps
  39. 5
  40. --speculative-eagle-topk
  41. 1
  42. --speculative-num-draft-tokens
  43. 6
  44. --speculative-adaptive
  45. --reasoning-parser
  46. glm45
  47. --tool-call-parser
  48. glm47
  49. --enable-multimodal
  50. --chat-template
  51. /chat-template.jinja
  52. --host
  53. 0.0.0.0
  54. --port
  55. 30000
FlagValue
--model-path/model
--served-model-nameglm-5.3-flash
--tp-size4
--ep-size4
--context-length1048576
--quantizationmodelopt_fp4
--attention-backenddsa
--dsa-prefill-backendflashinfer_sparse_mla
--dsa-decode-backendflashinfer_sparse_mla
--linear-attn-backendtriton
--kv-cache-dtypefp8_e4m3
--moe-runner-backendflashinfer_cutlass
--chunked-prefill-size8192
--max-prefill-tokens8192
--max-running-requests8
--mem-fraction-static0.93
--cuda-graph-max-bs-decode8
--speculative-algorithmNEXTN
--speculative-num-steps5
--speculative-eagle-topk1
--speculative-num-draft-tokens6
--reasoning-parserglm45
--tool-call-parserglm47
--chat-template/chat-template.jinja
--host0.0.0.0
--port30000

Environment

VariableValue
CUDA_DEVICE_ORDERPCI_BUS_ID
FLASHINFER_CUDA_ARCH_LIST12.0f
SGLANG_FP8_PAGED_MQA_LOGITS_TORCH1
SGLANG_OPT_DEEPGEMM_HC_PRENORM0
SGLANG_OPT_FP8_WO_A_GEMM0
SGLANG_OPT_USE_TILELANG_INDEXER1
SGLANG_OPT_USE_TOPK_V20
TORCH_CUDA_ARCH_LIST12.0a

Mounts

SourceTarget
${MODEL_ROOT}/GLM-5.3-Flash-NVFP4/model (read-only)

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1141.8124observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
2123252observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
4100.3256observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
872.5332observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
12,1255,344398observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
16,3256,383991observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
121,4076,2753,411observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
153,3156,1808,627observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
1106,7266,02917,702observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
1213,0945,78336,850observedglm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
glm53-flash-nvfp4-rtxpro6000-sglang-tp4
model instance id
libertaidai-glm-5-3-flash-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm53-flash-nvfp4-rtxpro6000-sglang-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
full decode capture bs<=8 with NEXTN target-verify graphs
name
sglang
version
lmsysorg/sglang:glm-5.3-flash@sha256:3a97bd50034ca60c6e6c86b8e36a73675d261f6a5eb71197796aee5175409290

serving

kv cache tokens
15,156,736
max concurrency
8
max context tokens
1,048,576
tensor parallel
4
Provenance & metadata (3)

facts

metadata

acceptance · completion
Yes
acceptance · tools
Yes
acceptance · video
Yes
acceptance · vision
Yes
runtime pin missing
launch.image digest

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry