Speed sweep

pg-01a011ee-c3fc-7cda-b3e2-f918f3126c25-sweep

pg-01a011ee-c3fc-7cda-b3e2-f918f3126c25-sweep

Record

Recipe
pg-bartowski-llama-3-1-nemotron-70b-instruct-hf-gguf-llama-3-1-6efbcc1dfea1-apple-m3-ultra-96gb-80c-llama-cpp-e8d1cfd50e
Measured
2026-08-17T22:54:15.543Z
Points
3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25613.5observed
132,83211.1357,890.8observed
1131,0086.3observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-01a011ee-c3fc-7cda-b3e2-f918f3126c25-sweep
measured at
2026-08-17T22:54:15.543Z
recipe id
pg-bartowski-llama-3-1-nemotron-70b-instruct-hf-gguf-llama-3-1-6efbcc1dfea1-apple-m3-ultra-96gb-80c-llama-cpp-e8d1cfd50e
schema version
local-ai-registry/v1

metrics

base memory bytes
85696561152
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
11.053
decode8k context tokens
8,256
decode8k tps
13.465
decode max context tokens
131,008
decode max context tps
6.329
decode mode
non-mtp
inference engine version
unknown
latest point at
2026-08-17T22:54:15.543Z
max context tokens
131,008
max prompt tokens
130,944
memory8k bytes
85789949952
memory8k context tokens
8,256
memory max context bytes
86345940992
memory max context tokens
131,008
peak generation tps
11.053
peak memory bytes
86345940992
peak prompt tps
91.559
point count
112
ttft32k cached prompt tokens
28,677
ttft32k context tokens
32,768
ttft32k seconds
357.891

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25613.465UnknownUnknown79.898Unknown1
132,83211.053UnknownUnknown80.416Unknown1
1131,0086.329UnknownUnknown80.416Unknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:01a011ee-c3fc-7cda-b3e2-f918f3126c25
repository
local.ai