Recipe
glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4
glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4Controller-backed candidate with recovered composed-checkpoint and local-image identities. Raw 524K-profile context and capability artifacts survive, but the current controller row is a later 400K/FP8 profile, so those older artifacts do not promote this configuration.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- controller-managed; captured 2026-08-23
- Graph
- full-and-piecewise
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 400,000
- Max concurrency
- 1
- KV cache tokens
- 672,296
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- yes
Hugging Face model card
Identity
https://huggingface.co/0xSero/GLM-5.2-TR3-Vision- Repository
- 0xSero/GLM-5.2-TR3-Vision
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Controller configuration
vllm
No container · candidate · controller
Candidate evidence — not a Run contract
Environment
| Variable | Value |
|---|---|
CUDA_DEVICE_ORDER | PCI_BUS_ID |
CUDA_VISIBLE_DEVICES | GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006 |
NVIDIA_VISIBLE_DEVICES | GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006 |
VLLM_B12X_MLA_CKV_PREFETCH_DEPTH | 0 |
VLLM_B12X_MLA_DCP_GATHER_IN_WORKSPACE | 0 |
VLLM_DCP_PROJECT_BEFORE_MERGE | 0 |
VLLM_DCP_QUERY_SPLIT | 0 |
VLLM_DCP_TOPK_OWNER_MERGE | 0 |
VLLM_ENABLE_PCIE_ALLREDUCE | 0 |
VLLM_USE_B12X_DCP_A2A | 0 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 8,192 | 1,804.3 | 21.1 | — | historical | glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4-sweep |
| 1 | 131,072 | 925.5 | 23.7 | — | historical | glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4-sweep |
| 1 | 400,000 | 852.5 | 16.9 | — | historical | glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4-sweep |
| 1 | — | — | 70.7 | — | historical | glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4
- model instance id
- 0xsero-glm-5-2-tr3-vision--3-0-bpw-text-bf16-vision
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm-52-vision-exl3-tr3-3bpw-vision-rtxpro6000-vllm-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- Yes
engine
- graph mode
- full-and-piecewise
- name
- vllm
- version
- controller-managed; captured 2026-08-23
serving
- kv cache tokens
- 672,296
- max concurrency
- 1
- max context tokens
- 400,000
- tensor parallel
- 4
Provenance & metadata (3)
facts
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |