Speed sweep
inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweep
inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweepRecord
- Measured
- 2026-08-26
- Points
- 5
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 30 | — | 117.8 | 60 | measured |
| 1 | 40 | — | 116 | 60 | measured |
| 1 | 40 | — | 115.6 | 63 | measured |
| 1 | 50 | — | 115.2 | 62 | measured |
| 4 | 40 | — | 84.2 | 320 | measured |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- id
- inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweep
- measured at
- 2026-08-26
- recipe id
- inkling-small-nvfp4-rtxpro6000-sglang-tp2
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 4
- inference engine version
- lmsysorg/sglang:dev-cu13-inkling-dspark@sha256:fbea1a4e25b26660dbc2384a27ead8817e9b7670f257b5c3143e0450d14524d7 + 1 patched file (SGLang commit b7252cc6b0c78b25ecea7ee5efa91a6ae37d0f19); built image local/sglang-inkling:sm120
- latest point at
- 2026-08-26
- max context tokens
- 50
- peak generation tps
- 306.6
- point count
- 5
rows
| concurrency | context tokens | context tokens note | decode tok s | decode tok s per stream | effective tok s incl ttft | output tokens | peak vram gb |
|---|---|---|---|---|---|---|---|
| 1 | 30 | estimated prompt length (bench JSON records no prompt_tokens) | 117.8 | 117.8 | 114.7 | 224 | Unknown |
| 1 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 116 | 116 | 115.6 | 1,528 | Unknown |
| 1 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 115.6 | 115.6 | 115.2 | 1,747 | Unknown |
| 1 | 50 | estimated prompt length (bench JSON records no prompt_tokens) | 115.2 | 115.2 | 115 | 2,500 | Unknown |
| 4 | 40 | estimated prompt length (bench JSON records no prompt_tokens) | 306.6 | 84.2 | 306.6 | 1,824 | Unknown |
Provenance & metadata (1)
source
- kind
- submitter
- methodology
- OpenAI /v1/chat/completions, streaming (SSE) with usage in the final chunk; temperature 0.7 / top_p 0.8; four prompt classes x 2 runs (short 400 cap, reasoning, analytical, long 2500 cap); effective tok/s = completion_tokens (reasoning + content, from usage) / wall incl. TTFT; TTFT = first streamed delta of content or reasoning; decode tok/s excludes TTFT; 4-way concurrency on the reasoning prompt, aggregate = total completion tokens / batch wall. Idle server, GPUs at the 300 W Max-Q limit. Raw JSON + script in the source repo benchmarks/.
- paths
- benchmarks/inkling-small-nvfp4-tp2-tps-20260826.json