Recipe

glm53-flash-exl3-q4-rtxpro6000-sglang-tp3

glm53-flash-exl3-q4-rtxpro6000-sglang-tp3

Three-GPU compatibility record for the selective EXL3 Q4 artifact. This is deliberately non-launchable: the published checkpoint is sealed as four independently rotated tensor-parallel slices and the loader requires TP4. Mapping four slices to three 96 GB GPUs would place two prepared slices, projected at 106.36 GB before KV cache and workspace, on one GPU.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
selective-EXL3 TP3 adaptation not implemented
Accelerators
3
Tensor parallel
3
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
Repository
0xSero/GLM-5.3-Flash-EXL3-Q4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

sglang

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4

No tokenized launch fields. This is measured or documented compatibility, not a Docker launch.

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1not-run-capacity-and-format-blockedglm53-flash-exl3-q4-rtxpro6000-sglang-tp3-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
3
hardware id
rtx-pro-6000-blackwell-96gb
id
glm53-flash-exl3-q4-rtxpro6000-sglang-tp3
model instance id
0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm53-flash-exl3-q4-rtxpro6000-sglang-tp3-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
python3, -m, sglang.launch_server, --model-path, 0xSero/GLM-5.3-Flash-EXL3-Q4, --tp, 3, --host, 0.0.0.0, --port, 30000
container port
30,000
host port
30,000
image
lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
gemma-4-12b-it-nvfp4-rtxpro6000-sglang-tp1
synthesized · template
sglang-launch-server-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
sglang
version
selective-EXL3 TP3 adaptation not implemented

serving

tensor parallel
3
Provenance & metadata (3)

facts

metadata

compatibility state
blocked
evidence · artifact bytes
187,376,244,600
evidence · artifact tensor parallel size
4
evidence · loader contract
TP=4 with EP=1 or EP=4
evidence · observed prepared weight gb per four gpu rank
53.18
evidence · projected double slice weight gb on one three gpu rank
106.36
required work
define a nonuniform four-logical-slice to three-physical-rank execution layout, preserve each independently rotated source slice without concatenation, fit the double-slice rank plus KV and workspace below 96 GB, pass real-weight parity and full CUDA-graph capture, run endpoint, completion, quality, and matched speed acceptance

provenance

captured at
2026-08-27T23:45:00Z

sources

captured atkindurl
2026-08-27T23:45:00Zartifact-contracthuggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4