Recipe
glm52-int4-int8mix-dgxspark-vllm-tp4
glm52-int4-int8mix-dgxspark-vllm-tp4Upstream four-Spark TP4/DCP4 candidate; dry-run structure is correct, but two upstream manifest hashes fail and the launcher expects an undocumented flattened overlay layout, so do not recommend or launch until those blockers are resolved
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- 0.27.0 + b12x 334a2d75d166becea0aa640b402d521ea0a290eb
- Graph
- full
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 315,968
- Max concurrency
- 16
- KV cache tokens
- 959,000
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/QuantTrio/GLM-5.2-Int4-Int8Mix- Repository
- QuantTrio/GLM-5.2-Int4-Int8Mix
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Script configuration
vllm
No container · candidate · script
Candidate evidence — not a Run contract
Environment
| Variable | Value |
|---|---|
DCP_SIZE | 4 |
GPU_MEMORY_UTILIZATION | 0.90 |
MAX_MODEL_LEN | 315968 |
MAX_NUM_BATCHED_TOKENS | 4096 |
MAX_NUM_SEQS | 16 |
MTP_SPEC_TOKENS | 2 |
TP_SIZE | 4 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | 618 | — | — | historical | glm52-int4-int8mix-dgxspark-vllm-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- dgx-spark-gb10-128gb
- id
- glm52-int4-int8mix-dgxspark-vllm-tp4
- model instance id
- quanttrio-glm-5-2-int4-int8mix--w4a16-int8mix
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm52-int4-int8mix-dgxspark-vllm-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- No
engine
- graph mode
- full
- name
- vllm
- version
- 0.27.0 + b12x 334a2d75d166becea0aa640b402d521ea0a290eb
serving
- kv cache tokens
- 959,000
- max concurrency
- 16
- max context tokens
- 315,968
- tensor parallel
- 4
Provenance & metadata (3)
facts
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |