Speed sweep

qwen3-30b-a3b-thinking-2507-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep

qwen3-30b-a3b-thinking-2507-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep

Record

Recipe
qwen3-30b-a3b-thinking-2507-q4-k-m-rtx-3060-12gb-llama-cpp-tp2
Measured
2026-08-18T13:15:42.190Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
165,5361,073.494.3observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
qwen3-30b-a3b-thinking-2507-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep
measured at
2026-08-18T13:15:42.190Z
recipe id
qwen3-30b-a3b-thinking-2507-q4-k-m-rtx-3060-12gb-llama-cpp-tp2
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
b10443
latest point at
2026-08-18T13:15:42.190Z
max context tokens
65,536
peak generation tps
94.29
peak prompt tps
1,073.4
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
165,53694.2994.29020.11,073.41
Provenance & metadata (1)