Speed sweep

openreasoning-nemotron-7b-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep

openreasoning-nemotron-7b-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep

Record

Recipe
openreasoning-nemotron-7b-q4-k-m-rtx-3060-12gb-llama-cpp-tp2
Measured
2026-07-12T04:51:09.737Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
15122,264.168245.3observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
openreasoning-nemotron-7b-q4-k-m-rtx-3060-12gb-llama-cpp-tp2-sweep
measured at
2026-07-12T04:51:09.737Z
recipe id
openreasoning-nemotron-7b-q4-k-m-rtx-3060-12gb-llama-cpp-tp2
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
13f2b28
latest point at
2026-07-12T04:51:09.737Z
max context tokens
512
peak generation tps
67.97
peak prompt tps
2,264.07
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
151267.9767.971287.32,264.071
Provenance & metadata (1)