Speed sweep
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweepRecord
- Measured
- 2026-08-27
- Accepted
- 2026-08-27
- Points
- 5
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 32 | — | 22.6 | — | accepted-c1-baseline |
| 8 | 32 | — | 14.7 | — | accepted-shared-prefix |
| 16 | 32 | — | 7.5 | — | accepted-first-cold-shared-prefix |
| 32 | 36 | — | 5.5 | — | accepted-cold-unique-prefix-throughput |
| 32 | 32 | — | 6.4 | — | accepted-warm-shared-prefix-throughput |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-08-27
- id
- glm53-flash-exl3-q4-rtxpro6000-sglang-pp3-sweep
- measured at
- 2026-08-27
- recipe id
- glm53-flash-exl3-q4-rtxpro6000-sglang-pp3
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 32
- inference engine version
- GLM-5.3 selective-EXL3 virtual-slice PP3 overlay
- latest point at
- 2026-08-27
- max context tokens
- 36
- peak generation tps
- 203.78
- point count
- 5
rows
| cache hit rate | concurrency | context tokens | decode tok s | decode tok s per stream | elapsed s | output tokens | peak vram gb |
|---|---|---|---|---|---|---|---|
| 0 | 1 | 32 | 22.596 | 22.596 | 11.33 | 256 | Unknown |
| Unknown | 8 | 32 | 117.369 | 14.671 | 17.449 | 256 | Unknown |
| Unknown | 16 | 32 | 120.735 | 7.546 | 33.925 | 256 | Unknown |
| 0 | 32 | 36 | 176.824 | 5.526 | 46.329 | 256 | 87.02 |
| Unknown | 32 | 32 | 203.78 | 6.368 | 40.2 | 256 | 87.02 |
Provenance & metadata (1)
source
- artifact revision
- 99cccdf0e8741715662c383828a9ea601990c125
- method
- OpenAI-compatible max_completion_tokens requests against the accepted TP1/PP3/EP1 runtime. Cold C32 used unique prompt prefixes and measured a zero cache-hit gauge; warm C32 reused the shared prefix. Client throughput is completed output tokens divided by batch wall time.