Speed sweep

qwen3-8-27b-ud-q4-k-xl-rtx-4090-24gb-llama-cpp-tp1-sweep

qwen3-8-27b-ud-q4-k-xl-rtx-4090-24gb-llama-cpp-tp1-sweep

Record

Recipe
qwen3-8-27b-ud-q4-k-xl-rtx-4090-24gb-llama-cpp-tp1
Measured
2026-08-19T05:07:28.598Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
1172,032968.591.6observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
qwen3-8-27b-ud-q4-k-xl-rtx-4090-24gb-llama-cpp-tp1-sweep
measured at
2026-08-19T05:07:28.598Z
recipe id
qwen3-8-27b-ud-q4-k-xl-rtx-4090-24gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
llama.cpp 7e4c0a96880dae4fc4268ad441f8a6446bd5460a (build 4)
latest point at
2026-08-19T05:07:28.598Z
max context tokens
172,032
peak generation tps
91.63
peak prompt tps
968.53
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
1172,03291.6391.6316,384Unknown968.531
Provenance & metadata (1)