Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-tp3
glm53-flash-exl3-q4-rtxpro6000-sglang-tp3Three-GPU compatibility record for the selective EXL3 Q4 artifact. This is deliberately non-launchable: the published checkpoint is sealed as four independently rotated tensor-parallel slices and the loader requires TP4. Mapping four slices to three 96 GB GPUs would place two prepared slices, projected at 106.36 GB before KV cache and workspace, on one GPU.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- sglang
- Engine version
- selective-EXL3 TP3 adaptation not implemented
- Accelerators
- 3
- Tensor parallel
- 3
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4- Repository
- 0xSero/GLM-5.3-Flash-EXL3-Q4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
sglang
Evidence only · candidate · reference
Candidate evidence — not a Run contract
No tokenized launch fields. This is measured or documented compatibility, not a Docker launch.
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | — | — | — | not-run-capacity-and-format-blocked | glm53-flash-exl3-q4-rtxpro6000-sglang-tp3-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 3
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp3
- model instance id
- 0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp3-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- python3, -m, sglang.launch_server, --model-path, 0xSero/GLM-5.3-Flash-EXL3-Q4, --tp, 3, --host, 0.0.0.0, --port, 30000
- container port
- 30,000
- host port
- 30,000
- image
- lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- gemma-4-12b-it-nvfp4-rtxpro6000-sglang-tp1
- synthesized · template
- sglang-launch-server-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- sglang
- version
- selective-EXL3 TP3 adaptation not implemented
serving
- tensor parallel
- 3
Provenance & metadata (3)
facts
metadata
- compatibility state
- blocked
- evidence · artifact bytes
- 187,376,244,600
- evidence · artifact tensor parallel size
- 4
- evidence · loader contract
- TP=4 with EP=1 or EP=4
- evidence · observed prepared weight gb per four gpu rank
- 53.18
- evidence · projected double slice weight gb on one three gpu rank
- 106.36
- required work
- define a nonuniform four-logical-slice to three-physical-rank execution layout, preserve each independently rotated source slice without concatenation, fit the double-slice rank plus KV and workspace below 96 GB, pass real-weight parity and full CUDA-graph capture, run endpoint, completion, quality, and matched speed acceptance
provenance
- captured at
- 2026-08-27T23:45:00Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T23:45:00Z | artifact-contract | huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4 ↗ |