Recipe

glm53-flash-exl3-q4-rtxpro6000-sglang-tp2

glm53-flash-exl3-q4-rtxpro6000-sglang-tp2

Two-GPU compatibility record for the selective EXL3 Q4 artifact. This is deliberately non-launchable: the published checkpoint is sealed for four tensor-parallel slices, the tested loader requires TP4, and a paired-slice TP2 projection exceeds the per-GPU prepared-weight budget before KV allocation. It exists to prevent a four-GPU command from being misrepresented as a working two-GPU recipe.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
selective-EXL3 TP2 adaptation not implemented
Accelerators
2
Tensor parallel
2
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4
Repository
0xSero/GLM-5.3-Flash-EXL3-Q4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

sglang

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4

No tokenized launch fields. This is measured or documented compatibility, not a Docker launch.

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1not-run-capacity-and-format-blockedglm53-flash-exl3-q4-rtxpro6000-sglang-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
rtx-pro-6000-blackwell-96gb
id
glm53-flash-exl3-q4-rtxpro6000-sglang-tp2
model instance id
0xsero-glm-5-3-flash-exl3-q4--selective-exl3-q4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
glm53-flash-exl3-q4-rtxpro6000-sglang-tp2-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
python3, -m, sglang.launch_server, --model-path, 0xSero/GLM-5.3-Flash-EXL3-Q4, --tp, 2, --host, 0.0.0.0, --port, 30000
container port
30,000
host port
30,000
image
lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
gemma-4-12b-it-nvfp4-rtxpro6000-sglang-tp1
synthesized · template
sglang-launch-server-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
sglang
version
selective-EXL3 TP2 adaptation not implemented

serving

tensor parallel
2
Provenance & metadata (3)

facts

draft launch.entrypoint · provenance · captured at
2026-09-01T01:41:26Z
draft launch.entrypoint · reason linux-amd64-container-config-entrypoint
draft launch.entrypoint · state known

metadata

compatibility state
blocked
evidence · artifact bytes
187,376,244,600
evidence · artifact tensor parallel size
4
evidence · observed prepared weight gb per four gpu rank
53.18
evidence · projected paired slice weight gb per two gpu rank
106.36
required work
pair sealed source ranks 0+1 and 2+3 without concatenating independent rotations, expand top-8 logical routes to top-16 virtual slice routes, pass real-weight kernel parity and full CUDA-graph capture, fit prepared weights plus KV below 96 GB per GPU, run endpoint, completion, quality, and matched speed acceptance

provenance

captured at
2026-08-27T22:05:00Z

sources

captured atkindurl
2026-08-27T22:05:00Zartifact-contracthuggingface.co/0xSero/GLM-5.3-Flash-EXL3-Q4