Recipe

glm52-nvfp4-aqlm-dgxspark-vllm-tp3

glm52-nvfp4-aqlm-dgxspark-vllm-tp3

MiaAI-Lab GLM-5.2 hybrid NVFP4+AQLM vision profile across three DGX Sparks with TP3, DCP1, MTP-3, and full CUDA graphs

Record

Status
candidate
Source
mialabs
Engine
vllm
Engine version
jarrelscy glm52-sm120 fork
Graph
full
Accelerators
3
Tensor parallel
3
Context tokens
40,000
Max concurrency
1
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid
Repository
jarrelscy/GLM-5.2-NVFP4-AQLM-hybrid
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Script configuration

vllm

No container · candidate · script

Candidate evidence — not a Run contract

Environment

VariableValue
DCP_SIZE1
ENABLE_MTP1
GPU_MEM_UTIL0.895
HF_REVISION53e0082eedebd806b63e19779c47905937d768ca
KV_CACHE_DTYPEnvfp4_ds_mla
KV_CACHE_MEMORY_BYTES11811160064
MAX_MODEL_LEN348160
MAX_NUM_BATCHED_TOKENS4096
MAX_NUM_SEQS1
MTP_SPEC_TOKENS3
TP_SIZE3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
121historicalglm52-nvfp4-aqlm-dgxspark-vllm-tp3-sweep
140,00013.7historicalglm52-nvfp4-aqlm-dgxspark-vllm-tp3-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
3
hardware id
dgx-spark-gb10-128gb
id
glm52-nvfp4-aqlm-dgxspark-vllm-tp3
model instance id
jarrelscy-glm-5-2-nvfp4-aqlm-hybrid--nvfp4-hot-experts-aqlm-2-bit-cold-experts
recipe source
mialabs
schema version
local-ai-registry/v1
speed sweep ids
glm52-nvfp4-aqlm-dgxspark-vllm-tp3-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
full
name
vllm
version
jarrelscy glm52-sm120 fork

serving

max concurrency
1
max context tokens
40,000
tensor parallel
3
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry