Speed sweep

pg-019fd7b6-2676-7ceb-b741-e49b0a4cacb7-sweep

pg-019fd7b6-2676-7ceb-b741-e49b0a4cacb7-sweep

Record

Recipe
pg-bartowski-llama-3-1-nemotron-70b-instruct-hf-gguf-llama-3-1-6efbcc1dfea1-apple-m5-pro-64gb-llama-cpp-3f4ccdc746
Measured
2026-08-06T15:34:26.677Z
Points
3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,2566.2observed
132,8324.9728,161.3observed
140,8964.6observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019fd7b6-2676-7ceb-b741-e49b0a4cacb7-sweep
measured at
2026-08-06T15:34:26.677Z
recipe id
pg-bartowski-llama-3-1-nemotron-70b-instruct-hf-gguf-llama-3-1-6efbcc1dfea1-apple-m5-pro-64gb-llama-cpp-3f4ccdc746
schema version
local-ai-registry/v1

metrics

base memory bytes
55797202944
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
4.9
decode8k context tokens
8,256
decode8k tps
6.216
decode max context tokens
40,896
decode max context tps
4.58
decode mode
non-mtp
inference engine version
unknown
latest point at
2026-08-06T15:34:26.677Z
max context tokens
40,896
max prompt tokens
40,832
memory8k bytes
55889936384
memory8k context tokens
8,256
memory max context bytes
56031019008
memory max context tokens
40,896
peak generation tps
4.9
peak memory bytes
56031019008
peak prompt tps
45.001
point count
35
ttft32k cached prompt tokens
28,677
ttft32k context tokens
32,768
ttft32k seconds
728.161

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,2566.216UnknownUnknown52.052Unknown1
132,8324.9UnknownUnknown52.183Unknown1
140,8964.58UnknownUnknown52.183Unknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019fd7b6-2676-7ceb-b741-e49b0a4cacb7
repository
local.ai