Speed sweep
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep
glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweepRecord
- Measured
- 2026-09-01
- Accepted
- 2026-09-01
- Points
- 1
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 8,182 | 1,764.9 | 112.1 | 4,636.4 | accepted-sustained |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-09-01
- id
- glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4-sweep
- measured at
- 2026-09-01
- recipe id
- glm-5-3-flash-nvfp4-dgx-spark-gb10-128gb-vllm-tp4
- schema version
- local-ai-registry/v1
metrics
- base memory bytes
- 358800470016
- base memory context tokens
- 0
- concurrency
- 1
- decode8k context tokens
- 8,182
- decode8k tps
- 112.118
- decode mode
- C1 structured JSON-schema matched-window with natural stop
- inference engine version
- 0.1.dev20051+g487ecf187
- latest point at
- 2026-09-01
- max context tokens
- 524,288
- max prompt tokens
- 8,183
- memory8k bytes
- 359068327936
- memory8k context tokens
- 8,182
- peak generation tps
- 117.992
- peak memory bytes
- 359068327936
- peak prompt tps
- 1,771.873
- point count
- 1
rows
| all stream window | base unified memory gib | cache state | concurrency | content class | context tokens | decode tok s | decode tok s per stream |
|---|---|---|---|---|---|---|---|
| first-emitted-token-through-last-emitted-token | 334.159 | cold-unique-cache-salt | 1 | structured-json-schema | 8,182 | 112.118 | 112.118 |
Provenance & metadata (1)
source
- kind
- live-runtime-capture
- paths
- docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-20260901.json, docs/notes/glm-5-3-flash-nvfp4-dgx-spark-vllm-tp4-compose.yml, registry/asset/glm-5-3-flash-multimodal-chat-template.jinja, registry/asset/glm-5-3-flash-vllm-sparse-attn-indexer-kpool.py
- repository
- github.com/0xSero/local-ai-registry ↗