Recipe

glm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4

glm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4

Pinned-source TP4/DCP4 GLM-5.2 REAP-469B NVFP4 candidate. The upstream bundle reports 250K context and graph capture, but its image is tag-only, the model revision is absent, and current acceptance has not been replayed.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
voipmonitor b12x build 20260608; exact commit unreported
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
250,000
Max concurrency
2
KV cache tokens
710,593
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.2-NVFP4-REAP-469B
Repository
0xSero/GLM-5.2-NVFP4-REAP-469B
Status
unknown
Link type
Exact Hub repository

hf api access unresolved

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0,1,2,3
NCCL_P2P_DISABLE1
VLLM_USE_B12X_MOE1
VLLM_USE_B12X_SPARSE_INDEXER1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
160historicalglm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4-sweep
164,0005,1004512,000historicalglm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
glm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4
model instance id
0xsero-glm-5-2-nvfp4-reap-469b--nvfp4-reap
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm52-nvfp4-reap-469b-rtxpro6000-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
voipmonitor b12x build 20260608; exact commit unreported

serving

kv cache tokens
710,593
max concurrency
2
max context tokens
250,000
tensor parallel
4
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry