Recipe

glm52-mxfp8-nvfp4-nf3-rtxpro6000-vllm-tp4

glm52-mxfp8-nvfp4-nf3-rtxpro6000-vllm-tp4

Immutable TP4/DCP4 deployment for the GLM-5.2 MXFP8/NVFP4/NF3 hybrid checkpoint. Recovered Pop evidence pins the image, checkpoint, vLLM, B12X, and FlashInfer and proves coherent long-context generation, but a current exact-config C1/C2/C4 sweep is still required before promotion.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
c382f1d28d5be2f867c216609408bdb424d6049a+b12x-e44cb77777a075790ebe9f7aa9f225d073aea109
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
180,007
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
Repository
madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0,1,2,3
CUTE_DSL_ARCHsm_120a
HYBRID_KEPTb12x_nf3
HYBRID_NF3b12x_nf3
VLLM_USE_B12X_MOE1
VLLM_USE_B12X_SPARSE_INDEXER1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1180,00758.2historicalglm52-mxfp8-nvfp4-nf3-rtxpro6000-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
glm52-mxfp8-nvfp4-nf3-rtxpro6000-vllm-tp4
model instance id
madeby561-glm-5-2-mxfp8-nvfp4-nf3-hybrid--mxfp8-nvfp4-nf3-hybrid
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm52-mxfp8-nvfp4-nf3-rtxpro6000-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
c382f1d28d5be2f867c216609408bdb424d6049a+b12x-e44cb77777a075790ebe9f7aa9f225d073aea109

serving

max context tokens
180,007
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max concurrency · reason not-observed
serving.max concurrency · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry