Speed sweep
deepseek-v4-flash-0731-fp8-fp4-rtxpro6000-sglang-tp2-sweep
deepseek-v4-flash-0731-fp8-fp4-rtxpro6000-sglang-tp2-sweepRecord
- Measured
- 2026-08-26
- Points
- 5
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 30 | — | 57.9 | 80 | measured |
| 1 | 40 | — | 57.8 | 84 | measured |
| 1 | 40 | — | 57.9 | 80 | measured |
| 1 | 50 | — | 57.8 | 80 | measured |
| 4 | 40 | — | 47.2 | 364 | measured |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- id
- deepseek-v4-flash-0731-fp8-fp4-rtxpro6000-sglang-tp2-sweep
- measured at
- 2026-08-26
- recipe id
- deepseek-v4-flash-0731-fp8-fp4-rtxpro6000-sglang-tp2
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 4
- latest point at
- 2026-08-26
- max context tokens
- 50
- peak generation tps
- 155.2
- point count
- 5
rows
| concurrency | context tokens | context tokens note | decode tok s | decode tok s per stream | effective tok s incl ttft | output tokens | peak vram gb |
|---|---|---|---|---|---|---|---|
| 1 | 30 | estimated prompt length (bench JSON records no prompt_tokens) | 57.85 | 57.85 | 56.15 | 118 | Unknown |
| 1 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 57.8 | 57.8 | 57.6 | 987 | Unknown |
| 1 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 57.85 | 57.85 | 57.7 | 1,995 | Unknown |
| 1 | 50 | estimated prompt length (bench JSON records no prompt_tokens) | 57.75 | 57.75 | 57.7 | 2,500 | Unknown |
| 4 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 155.2 | 47.2 | 155.2 | 1,611 | Unknown |
Provenance & metadata (1)
source
- kind
- submitter
- methodology
- OpenAI /v1/chat/completions, streaming (SSE) with usage in the final chunk; temperature 0.7 / top_p 0.8; four prompt classes x 2 runs (short 400 cap, reasoning, analytical, long 2500 cap); effective tok/s = completion_tokens (reasoning + content, from usage) / wall incl. TTFT; TTFT = first streamed delta of content or reasoning; decode tok/s excludes TTFT; 4-way concurrency on the reasoning prompt, aggregate = total completion tokens / batch wall. Idle server, GPUs at the 300 W Max-Q limit. Raw JSON + script in the source repo benchmarks/.
- paths
- benchmarks/deepseek-v4-flash-0731-tp2-tps-20260826.json