Speed sweep

llama-3-2-3b-instruct-q4-k-m-rtx-3060-12gb-llama-cpp-tp1-sweep

llama-3-2-3b-instruct-q4-k-m-rtx-3060-12gb-llama-cpp-tp1-sweep

Record

Recipe
llama-3-2-3b-instruct-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
Measured
2026-07-24T03:51:56.228Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
15124,321.8128.1174.5observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
llama-3-2-3b-instruct-q4-k-m-rtx-3060-12gb-llama-cpp-tp1-sweep
measured at
2026-07-24T03:51:56.228Z
recipe id
llama-3-2-3b-instruct-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
13f2b28
latest point at
2026-07-24T03:51:56.228Z
max context tokens
512
peak generation tps
128.08
peak prompt tps
4,321.8
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
1512128.08128.08128Unknown4,321.81
Provenance & metadata (1)