Recipe
glm-5.2-nvfp4-b200-vllm-tp8-mtp5
glm-5.2-nvfp4-b200-vllm-tp8-mtp5Local AI PostgreSQL candidate reconstructed from two completed GLM-5.2 evaluations on eight B200 GPUs. Model and image are now immutably pinned, but the source did not preserve graph-mode, Docker IPC/shared-memory settings, an on-hardware completion artifact, or a speed sweep, so CLI launch remains blocked.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- 0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
- Graph
- unknown
- Accelerators
- 8
- Tensor parallel
- 8
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/nvidia/GLM-5.2-NVFP4- Repository
- nvidia/GLM-5.2-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
vllm
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/davidmcc73/vllm-openai@sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2- Digest
sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2- Port
- 8000
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--modelnvidia/GLM-5.2-NVFP4--served-model-namenvidia/GLM-5.2-NVFP4--host0.0.0.0--port8000--tensor-parallel-size8--pipeline-parallel-size1--trust-remote-code--enable-chunked-prefill--enable-prefix-caching--enable-auto-tool-choice--enable-prompt-tokens-details--enable-force-include-usage--enable-request-id-headers--enable-log-requests--max-num-seqs1--gpu-memory-utilization0.90--block-size64--language-model-only--enable-expert-parallel--max-model-len1048576--max-num-batched-tokens8192--tool-call-parserglm47--reasoning-parserglm45--kv-cache-dtypefp8_e4m3--speculative-config{"method":"mtp","num_speculative_tokens":5,"rejection_sample_method":"standard"}
| Flag | Value |
|---|---|
--model | nvidia/GLM-5.2-NVFP4 |
--served-model-name | nvidia/GLM-5.2-NVFP4 |
--host | 0.0.0.0 |
--port | 8000 |
--tensor-parallel-size | 8 |
--pipeline-parallel-size | 1 |
--max-num-seqs | 1 |
--gpu-memory-utilization | 0.90 |
--block-size | 64 |
--max-model-len | 1048576 |
--max-num-batched-tokens | 8192 |
--tool-call-parser | glm47 |
--reasoning-parser | glm45 |
--kv-cache-dtype | fp8_e4m3 |
--speculative-config | {"method":"mtp","num_speculative_tokens":5,"rejection_sample_method":"standard"} |
Environment
| Variable | Value |
|---|---|
HF_HOME | /root/.cache/huggingface |
Mounts
| Source | Target |
|---|---|
~/.cache/huggingface | /root/.cache/huggingface |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 8
- hardware id
- b200-180gb
- id
- glm-5.2-nvfp4-b200-vllm-tp8-mtp5
- model instance id
- nvidia-glm-5-2-nvfp4--nvfp4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- No
engine
- graph mode
- unknown
- name
- vllm
- version
- 0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
serving
- tensor parallel
- 8
Provenance & metadata (3)
facts
- serving.kv cache tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max concurrency · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max context tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- speed sweep ids · provenance · captured at
- 2026-08-27T06:04:13.773Z
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |