Recipe

glm52-nvfp4-reap-nu176-rtxpro6000-vllm-tp4

glm52-nvfp4-reap-nu176-rtxpro6000-vllm-tp4

Pinned-source non-uniform REAP NU176 candidate for four RTX PRO 6000 Blackwell GPUs. The source reports 262K context, concurrency two, MTP4, and full/piecewise CUDA graphs, but model and image digest pins plus current index acceptance are still missing.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
b12x nonuniform fork 20260705; exact commit unreported
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
262,144
Max concurrency
2
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.2-REAP-NU176-526B
Repository
0xSero/GLM-5.2-REAP-NU176-526B
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0,1,2,3
NCCL_P2P_DISABLE1
VLLM_USE_B12X_MOE1
VLLM_USE_B12X_SPARSE_INDEXER1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
177historicalglm52-nvfp4-reap-nu176-rtxpro6000-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
glm52-nvfp4-reap-nu176-rtxpro6000-vllm-tp4
model instance id
0xsero-glm-5-2-reap-nu176-526b--nvfp4-non-uniform-reap
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm52-nvfp4-reap-nu176-rtxpro6000-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
b12x nonuniform fork 20260705; exact commit unreported

serving

max concurrency
2
max context tokens
262,144
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry