Speed sweep

glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep

glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep

Record

Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
Measured
2026-08-27
Accepted
2026-08-27
Points
5

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
13222.6accepted-c1-baseline
83214.7accepted-shared-prefix
16327.5accepted-first-cold-shared-prefix
32365.5accepted-cold-unique-prefix-throughput
32326.4accepted-warm-shared-prefix-throughput

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

accepted at
2026-08-27
id
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
measured at
2026-08-27
recipe id
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
schema version
local-ai-registry/v1

metrics

concurrency
32
inference engine version
GLM-5.3 selective-EXL3 virtual-slice PP3 overlay
latest point at
2026-08-27
max context tokens
36
peak generation tps
203.78
point count
5

rows

cache hit rateconcurrencycontext tokensdecode tok sdecode tok s per streamelapsed soutput tokenspeak vram gb
013222.59622.59611.33256Unknown
Unknown832117.36914.67117.449256Unknown
Unknown1632120.7357.54633.925256Unknown
03236176.8245.52646.32925687.02
Unknown3232203.786.36840.225687.02
Provenance & metadata (1)

source

artifact revision
99cccdf0e8741715662c383828a9ea601990c125
method
OpenAI-compatible max_completion_tokens requests against the accepted TP1/PP3/EP1 runtime. Cold C32 used unique prompt prefixes and measured a zero cache-hit gauge; warm C32 reused the shared prefix. Client throughput is completed output tokens divided by batch wall time.