Speed sweep
glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep
glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweepRecord
- Measured
- 2026-09-01T03:45:33.045685Z
- Accepted
- 2026-09-01T04:00:19Z
- Points
- 3
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 78 | 1,752.2 | 61.2 | 325.5 | validated |
| 2 | 78 | 1,768.1 | 42.9 | 503.1 | validated |
| 3 | 78 | 1,855.5 | 38.5 | 675 | validated |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-09-01T04:00:19Z
- id
- glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep
- measured at
- 2026-09-01T03:45:33.045685Z
- recipe id
- glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-96gb-vllm-tp4
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 3
- decode mode
- sustained-matched-window
- inference engine version
- 0.26.1rc0+infernal.invocation.cu133.r17.vllmc53cc73.b12xc0a44a1
- latest point at
- 2026-09-01T03:45:33.045685Z
- max context tokens
- 800,000
- max prompt tokens
- 8,192
- peak generation tps
- 115.542
- peak memory bytes
- 397263503360
- peak prompt tps
- 1,855.53
- point count
- 3
rows
| aggregate source | capacity limited | concurrency | context tokens | decode tok s | decode tok s per stream | effective concurrency | errors |
|---|---|---|---|---|---|---|---|
| openai_continuous_usage | No | 1 | 78 | 61.211 | 61.211 | 1 | 0 |
| openai_continuous_usage | No | 2 | 78 | 85.783 | 42.891 | 2 | 0 |
| openai_continuous_usage | No | 3 | 78 | 115.542 | 38.514 | 3 | 0 |
Provenance & metadata (3)
facts
- metrics.peak generation tps · provenance · captured at
- 2026-09-01T04:00:19Z
metrics.peak generation tps · reason openai-continuous-usage-completion-tokens-divided-by-common-measurement-window
metrics.peak generation tps · state known
- metrics.peak memory bytes · provenance · captured at
- 2026-09-01T04:00:19Z
metrics.peak memory bytes · reason maximum-summed-memory-used-from-340-half-second-four-gpu-samples
metrics.peak memory bytes · state known
- metrics.peak prompt tps · provenance · captured at
- 2026-09-01T04:00:19Z
metrics.peak prompt tps · reason exact-prompt-tokens-divided-by-common-wall-time-to-first-token-with-uncapped-natural-stop-requests
metrics.peak prompt tps · state known
provenance
- captured at
- 2026-09-01T04:00:19Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-09-01T04:00:19Z | validated-speed-sweep | github.com/0xSero/local-ai-registry/blob/main/docs/notes/glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-20260901.json ↗ |
source
- kind
- live-gpu-benchmark
- paths
- docs/notes/glm-5-3-exl3-tr3-3-0bpw-rtx-pro-6000-blackwell-20260901.json
- repository
- github.com/0xSero/local-ai-registry ↗