Speed sweep
deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweepRecord
- Measured
- 2026-08-24
- Accepted
- 2026-08-24
- Points
- 7
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 8,192 | 701.6 | 37.6 | 11,337 | historical |
| 1 | 16,384 | 1,250 | 38 | 12,770 | historical |
| 1 | 32,768 | 1,273.1 | 35.5 | 25,101 | historical |
| 1 | 65,536 | 1,145.1 | 30.3 | 55,850 | historical |
| 1 | 131,072 | 946 | 24.1 | 135,271 | historical |
| 1 | 196,000 | 793.8 | 19.3 | 246,863 | historical |
| 1 | 300,000 | 646.4 | 15.1 | 464,072 | historical |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-08-24
- id
- deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
- measured at
- 2026-08-24
- recipe id
- deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 1
- latest point at
- 2026-08-24
- max context tokens
- 300,000
- peak generation tps
- 37.99
- peak prompt tps
- 1,273.1
- point count
- 7
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| 1 | 8,192 | 37.59 | 37.59 | Unknown | Unknown | 701.6 | 1 |
| 1 | 16,384 | 37.99 | 37.99 | Unknown | Unknown | 1,250 | 1 |
| 1 | 32,768 | 35.5 | 35.5 | Unknown | Unknown | 1,273.1 | 1 |
| 1 | 65,536 | 30.33 | 30.33 | Unknown | Unknown | 1,145.1 | 1 |
| 1 | 131,072 | 24.08 | 24.08 | Unknown | Unknown | 946 | 1 |
| 1 | 196,000 | 19.33 | 19.33 | Unknown | Unknown | 793.8 | 1 |
| 1 | 300,000 | 15.09 | 15.09 | Unknown | Unknown | 646.4 | 1 |
Provenance & metadata (1)
source
- commit
- c2eac5a9b2b457881d69b1164d909e8beab9286e
- paths
- README.md, DEEPSEEK-V4-FLASH.md, scripts/launch_dsv4_flash_sm120.sh, scripts/build_in_sglang_docker.sh, deepseek_v4_kernel/_patch.py