Speed sweep

glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep

glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep

Record

Recipe
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
Measured
2026-08-27
Accepted
2026-08-27
Points
19

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
12973.5226accepted-matched-screen
12968.5167.1rejected-matched-regression
12966observed-earlier-tp4-ep1-baseline
12964.9rejected-speculation-regression
11,024827.71,248accepted-cold-prefill
11,022140.3466.8accepted-warm-prefix-prefill
1126550.579238.1accepted-cold-c1
2128518.563326.6accepted-cold-concurrency-sweep
4128295.962.3671.6accepted-cold-concurrency-sweep
8128378.355.4699.6accepted-cold-concurrency-sweep
16128327.838.21,538.6accepted-cold-concurrency-sweep
3212834330.42,299.7accepted-balanced-latency-profile
6412836120.93,732.3accepted-short-output-throughput-profile
128128371.417.319,334.7rejected-short-output-throughput-regression
64131191.928.92,800accepted-sustained-warm-prefix
128131211.628.438,464.4accepted-sustained-warm-prefix
256131222.631.668,677.4accepted-max-output-tpm-profile
64127295.817.37,432.7rejected-ep4-saturation-regression
6412734520.13,617rejected-single-batch-overlap-regression

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

accepted at
2026-08-27
id
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
measured at
2026-08-27
recipe id
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
schema version
local-ai-registry/v1

metrics

concurrency
256
inference engine version
0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay
latest point at
2026-08-27
max context tokens
1,024
peak generation tps
1,710.164
peak prompt tps
827.74
point count
19

rows

concurrencycontext tokensdecode tok sdecode tok s per streamelapsed sengine decode tok sexpert paralleloutput tokens
12973.49873.4983.48397.0521256
12968.46668.4663.73988.5644256
12965.96665.9663.881UnknownUnknown256
12964.86164.8613.947UnknownUnknown256
11,024UnknownUnknown1.248UnknownUnknown1
11,022UnknownUnknown0.467UnknownUnknown1
112673.61979.0343.477UnknownUnknown256
2128116.66663.0294.389UnknownUnknown256
4128214.04962.264.784UnknownUnknown256
8128384.95155.4195.32UnknownUnknown256
16128496.98238.28.242UnknownUnknown256
32128763.7430.39510.726858.654Unknown256
64128812.02620.90720.177UnknownUnknown256
128128798.05217.25641.06UnknownUnknown256
641311,183.65828.90155.367UnknownUnknown1,024
1281311,423.86228.38192.054UnknownUnknown1,024
2561311,710.16431.607153.286UnknownUnknown1,024
64127563.98117.26729.051Unknown4256
64127792.00920.07420.687UnknownUnknown256
Provenance & metadata (1)

source

artifact revision
99cccdf0e8741715662c383828a9ea601990c125
method
OpenAI-compatible streaming sweeps with server-reported prompt/completion counts. Cold ISL/OSL concurrency screens use distinct prompts; sustained screens prime one shared prefix and request OSL 1024. Client output rate is total completed output tokens divided by batch wall time. Output TPM is output rate multiplied by 60; total API TPM includes prompt tokens. SGLang Prometheus deltas supply prefill-compute tokens, device-cache hits, and engine gauges. Full CUDA graphs remained enabled throughout.
runtime image
glm53-flash-sglang-exl3:stage-20260827-r6