Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-tp2
glm53-flash-exl3-q4-rtxpro6000-sglang-tp2Two-GPU compatibility record for the selective EXL3 Q4 artifact. This is deliberately non-launchable: the published checkpoint is sealed for four tensor-parallel slices, the tested loader requires TP4, and a paired-slice TP2 projection exceeds the per-GPU prepared-weight budget before KV allocation. It exists to prevent a four-GPU command from being misrepresented as a working two-GPU recipe.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- sglang
- Engine version
- selective-EXL3 TP2 adaptation not implemented
- Accelerators
- 2
- Tensor parallel
- 2
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4- Repository
- 0xSero/GLM-5.3-Flash-EXL3-Q4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
sglang
Evidence only · candidate · reference
Candidate evidence — not a Run contract
No tokenized launch fields. This is measured or documented compatibility, not a Docker launch.
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | — | — | — | not-run-capacity-and-format-blocked | glm53-flash-exl3-q4-rtxpro6000-sglang-tp2-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 2
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp2
- model instance id
- 0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp2-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- python3, -m, sglang.launch_server, --model-path, 0xSero/GLM-5.3-Flash-EXL3-Q4, --tp, 2, --host, 0.0.0.0, --port, 30000
- container port
- 30,000
- entrypoint
- /opt/nvidia/nvidia_entrypoint.sh
- host port
- 30,000
- image
- lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- gemma-4-12b-it-nvfp4-rtxpro6000-sglang-tp1
- synthesized · template
- sglang-launch-server-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- sglang
- version
- selective-EXL3 TP2 adaptation not implemented
serving
- tensor parallel
- 2
Provenance & metadata (3)
facts
- draft launch.entrypoint · provenance · captured at
- 2026-09-01T01:41:26Z
metadata
- compatibility state
- blocked
- evidence · artifact bytes
- 187,376,244,600
- evidence · artifact tensor parallel size
- 4
- evidence · observed prepared weight gb per four gpu rank
- 53.18
- evidence · projected paired slice weight gb per two gpu rank
- 106.36
- required work
- pair sealed source ranks 0+1 and 2+3 without concatenating independent rotations, expand top-8 logical routes to top-16 virtual slice routes, pass real-weight kernel parity and full CUDA-graph capture, fit prepared weights plus KV below 96 GB per GPU, run endpoint, completion, quality, and matched speed acceptance
provenance
- captured at
- 2026-08-27T22:05:00Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T22:05:00Z | artifact-contract | huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4 ↗ |