Speed sweep

nvidia-nemotron-labs-3-elastic-30b-a3b-nvfp4-q4-k-s-rtx-5060-ti-16gb-llama-cpp-tp1-sweep

nvidia-nemotron-labs-3-elastic-30b-a3b-nvfp4-q4-k-s-rtx-5060-ti-16gb-llama-cpp-tp1-sweep

Record

Recipe
nvidia-nemotron-labs-3-elastic-30b-a3b-nvfp4-q4-k-s-rtx-5060-ti-16gb-llama-cpp-tp1
Measured
2026-05-15T02:01:01.280Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
132,768694.4158.369observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
nvidia-nemotron-labs-3-elastic-30b-a3b-nvfp4-q4-k-s-rtx-5060-ti-16gb-llama-cpp-tp1-sweep
measured at
2026-05-15T02:01:01.280Z
recipe id
nvidia-nemotron-labs-3-elastic-30b-a3b-nvfp4-q4-k-s-rtx-5060-ti-16gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
latest point at
2026-05-15T02:01:01.280Z
max context tokens
32,768
peak generation tps
158.28
peak prompt tps
694.38
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
132,768158.28Unknown1,02415694.381
Provenance & metadata (1)