Speed sweep

inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweep

inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweep

Record

Recipe
inkling-small-nvfp4-rtxpro6000-sglang-tp2
Measured
2026-08-26
Points
5

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
130117.860measured
14011660measured
140115.663measured
150115.262measured
44084.2320measured

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
inkling-small-nvfp4-rtxpro6000-sglang-tp2-sweep
measured at
2026-08-26
recipe id
inkling-small-nvfp4-rtxpro6000-sglang-tp2
schema version
local-ai-registry/v1

metrics

concurrency
4
inference engine version
lmsysorg/sglang:dev-cu13-inkling-dspark@sha256:fbea1a4e25b26660dbc2384a27ead8817e9b7670f257b5c3143e0450d14524d7 + 1 patched file (SGLang commit b7252cc6b0c78b25ecea7ee5efa91a6ae37d0f19); built image local/sglang-inkling:sm120
latest point at
2026-08-26
max context tokens
50
peak generation tps
306.6
point count
5

rows

concurrencycontext tokenscontext tokens notedecode tok sdecode tok s per streameffective tok s incl ttftoutput tokenspeak vram gb
130estimated prompt length (bench JSON records no prompt_tokens)117.8117.8114.7224Unknown
140estimated prompt length (bench JSON records no prompt_tokens)116116115.61,528Unknown
140estimated prompt length (bench JSON records no prompt_tokens)115.6115.6115.21,747Unknown
150estimated prompt length (bench JSON records no prompt_tokens)115.2115.21152,500Unknown
440estimated prompt length (bench JSON records no prompt_tokens)306.684.2306.61,824Unknown
Provenance & metadata (1)

source

kind
submitter
methodology
OpenAI /v1/chat/completions, streaming (SSE) with usage in the final chunk; temperature 0.7 / top_p 0.8; four prompt classes x 2 runs (short 400 cap, reasoning, analytical, long 2500 cap); effective tok/s = completion_tokens (reasoning + content, from usage) / wall incl. TTFT; TTFT = first streamed delta of content or reasoning; decode tok/s excludes TTFT; 4-way concurrency on the reasoning prompt, aggregate = total completion tokens / batch wall. Idle server, GPUs at the 300 W Max-Q limit. Raw JSON + script in the source repo benchmarks/.
paths
benchmarks/inkling-small-nvfp4-tp2-tps-20260826.json