Speed sweep
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweepRecord
- Measured
- 2026-08-27
- Accepted
- 2026-08-27
- Points
- 19
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 29 | — | 73.5 | 226 | accepted-matched-screen |
| 1 | 29 | — | 68.5 | 167.1 | rejected-matched-regression |
| 1 | 29 | — | 66 | — | observed-earlier-tp4-ep1-baseline |
| 1 | 29 | — | 64.9 | — | rejected-speculation-regression |
| 1 | 1,024 | 827.7 | — | 1,248 | accepted-cold-prefill |
| 1 | 1,022 | 140.3 | — | 466.8 | accepted-warm-prefix-prefill |
| 1 | 126 | 550.5 | 79 | 238.1 | accepted-cold-c1 |
| 2 | 128 | 518.5 | 63 | 326.6 | accepted-cold-concurrency-sweep |
| 4 | 128 | 295.9 | 62.3 | 671.6 | accepted-cold-concurrency-sweep |
| 8 | 128 | 378.3 | 55.4 | 699.6 | accepted-cold-concurrency-sweep |
| 16 | 128 | 327.8 | 38.2 | 1,538.6 | accepted-cold-concurrency-sweep |
| 32 | 128 | 343 | 30.4 | 2,299.7 | accepted-balanced-latency-profile |
| 64 | 128 | 361 | 20.9 | 3,732.3 | accepted-short-output-throughput-profile |
| 128 | 128 | 371.4 | 17.3 | 19,334.7 | rejected-short-output-throughput-regression |
| 64 | 131 | 191.9 | 28.9 | 2,800 | accepted-sustained-warm-prefix |
| 128 | 131 | 211.6 | 28.4 | 38,464.4 | accepted-sustained-warm-prefix |
| 256 | 131 | 222.6 | 31.6 | 68,677.4 | accepted-max-output-tpm-profile |
| 64 | 127 | 295.8 | 17.3 | 7,432.7 | rejected-ep4-saturation-regression |
| 64 | 127 | 345 | 20.1 | 3,617 | rejected-single-batch-overlap-regression |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-08-27
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp4-sweep
- measured at
- 2026-08-27
- recipe id
- glm53-flash-exl3-q4-rtxpro6000-sglang-tp4
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 256
- inference engine version
- 0.0.0.dev1+gf609d677b with Sparkinfer 1.0.1 selective-EXL3 overlay
- latest point at
- 2026-08-27
- max context tokens
- 1,024
- peak generation tps
- 1,710.164
- peak prompt tps
- 827.74
- point count
- 19
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | elapsed s | engine decode tok s | expert parallel | output tokens |
|---|---|---|---|---|---|---|---|
| 1 | 29 | 73.498 | 73.498 | 3.483 | 97.052 | 1 | 256 |
| 1 | 29 | 68.466 | 68.466 | 3.739 | 88.564 | 4 | 256 |
| 1 | 29 | 65.966 | 65.966 | 3.881 | Unknown | Unknown | 256 |
| 1 | 29 | 64.861 | 64.861 | 3.947 | Unknown | Unknown | 256 |
| 1 | 1,024 | Unknown | Unknown | 1.248 | Unknown | Unknown | 1 |
| 1 | 1,022 | Unknown | Unknown | 0.467 | Unknown | Unknown | 1 |
| 1 | 126 | 73.619 | 79.034 | 3.477 | Unknown | Unknown | 256 |
| 2 | 128 | 116.666 | 63.029 | 4.389 | Unknown | Unknown | 256 |
| 4 | 128 | 214.049 | 62.26 | 4.784 | Unknown | Unknown | 256 |
| 8 | 128 | 384.951 | 55.419 | 5.32 | Unknown | Unknown | 256 |
| 16 | 128 | 496.982 | 38.2 | 8.242 | Unknown | Unknown | 256 |
| 32 | 128 | 763.74 | 30.395 | 10.726 | 858.654 | Unknown | 256 |
| 64 | 128 | 812.026 | 20.907 | 20.177 | Unknown | Unknown | 256 |
| 128 | 128 | 798.052 | 17.256 | 41.06 | Unknown | Unknown | 256 |
| 64 | 131 | 1,183.658 | 28.901 | 55.367 | Unknown | Unknown | 1,024 |
| 128 | 131 | 1,423.862 | 28.381 | 92.054 | Unknown | Unknown | 1,024 |
| 256 | 131 | 1,710.164 | 31.607 | 153.286 | Unknown | Unknown | 1,024 |
| 64 | 127 | 563.981 | 17.267 | 29.051 | Unknown | 4 | 256 |
| 64 | 127 | 792.009 | 20.074 | 20.687 | Unknown | Unknown | 256 |
Provenance & metadata (1)
source
- artifact revision
- 99cccdf0e8741715662c383828a9ea601990c125
- method
- OpenAI-compatible streaming sweeps with server-reported prompt/completion counts. Cold ISL/OSL concurrency screens use distinct prompts; sustained screens prime one shared prefix and request OSL 1024. Client output rate is total completed output tokens divided by batch wall time. Output TPM is output rate multiplied by 60; total API TPM includes prompt tokens. SGLang Prometheus deltas supply prefill-compute tokens, device-cache hits, and engine gauges. Full CUDA graphs remained enabled throughout.
- runtime image
- glm53-flash-sglang-exl3:stage-20260827-r6