Recipe

glm-5.2-nvfp4-b200-vllm-tp8-mtp5

glm-5.2-nvfp4-b200-vllm-tp8-mtp5

Local AI PostgreSQL candidate reconstructed from two completed GLM-5.2 evaluations on eight B200 GPUs. Model and image are now immutably pinned, but the source did not preserve graph-mode, Docker IPC/shared-memory settings, an on-hardware completion artifact, or a speed sweep, so CLI launch remains blocked.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
Graph
unknown
Accelerators
8
Tensor parallel
8
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/nvidia/GLM-5.2-NVFP4
Repository
nvidia/GLM-5.2-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

vllm

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/davidmcc73/vllm-openai@sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2
Digest
sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2
Port
8000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model
  2. nvidia/GLM-5.2-NVFP4
  3. --served-model-name
  4. nvidia/GLM-5.2-NVFP4
  5. --host
  6. 0.0.0.0
  7. --port
  8. 8000
  9. --tensor-parallel-size
  10. 8
  11. --pipeline-parallel-size
  12. 1
  13. --trust-remote-code
  14. --enable-chunked-prefill
  15. --enable-prefix-caching
  16. --enable-auto-tool-choice
  17. --enable-prompt-tokens-details
  18. --enable-force-include-usage
  19. --enable-request-id-headers
  20. --enable-log-requests
  21. --max-num-seqs
  22. 1
  23. --gpu-memory-utilization
  24. 0.90
  25. --block-size
  26. 64
  27. --language-model-only
  28. --enable-expert-parallel
  29. --max-model-len
  30. 1048576
  31. --max-num-batched-tokens
  32. 8192
  33. --tool-call-parser
  34. glm47
  35. --reasoning-parser
  36. glm45
  37. --kv-cache-dtype
  38. fp8_e4m3
  39. --speculative-config
  40. {"method":"mtp","num_speculative_tokens":5,"rejection_sample_method":"standard"}
FlagValue
--modelnvidia/GLM-5.2-NVFP4
--served-model-namenvidia/GLM-5.2-NVFP4
--host0.0.0.0
--port8000
--tensor-parallel-size8
--pipeline-parallel-size1
--max-num-seqs1
--gpu-memory-utilization0.90
--block-size64
--max-model-len1048576
--max-num-batched-tokens8192
--tool-call-parserglm47
--reasoning-parserglm45
--kv-cache-dtypefp8_e4m3
--speculative-config{"method":"mtp","num_speculative_tokens":5,"rejection_sample_method":"standard"}

Environment

VariableValue
HF_HOME/root/.cache/huggingface

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
8
hardware id
b200-180gb
id
glm-5.2-nvfp4-b200-vllm-tp8-mtp5
model instance id
nvidia-glm-5-2-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
unknown
name
vllm
version
0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96

serving

tensor parallel
8
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max concurrency · reason not-observed
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
speed sweep ids · provenance · captured at
2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry