Speed sweep

pg-01a011ee-e409-783c-a752-119b6ceea579-sweep

pg-01a011ee-e409-783c-a752-119b6ceea579-sweep

Record

Recipe
pg-unsloth-nvidia-nemotron-3-super-120b-a12b-gguf-ud-q3-k-m-apple-m4-max-128gb-llama-cpp-b3273677d9
Measured
2026-08-17T22:54:23.748Z
Points
3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25628.8observed
132,83227.895,269.9observed
1262,08019.8observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-01a011ee-e409-783c-a752-119b6ceea579-sweep
measured at
2026-08-17T22:54:23.748Z
recipe id
pg-unsloth-nvidia-nemotron-3-super-120b-a12b-gguf-ud-q3-k-m-apple-m4-max-128gb-llama-cpp-b3273677d9
schema version
local-ai-registry/v1

metrics

base memory bytes
63143772160
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
27.794
decode8k context tokens
8,256
decode8k tps
28.833
decode max context tokens
262,080
decode max context tps
19.814
decode mode
non-mtp
inference engine version
built with AppleClang 17.0.0.17000603 for Darwin arm64
latest point at
2026-08-17T22:54:23.748Z
max context tokens
262,080
max prompt tokens
262,016
memory8k bytes
64418299904
memory8k context tokens
8,256
memory max context bytes
69884854272
memory max context tokens
262,080
peak generation tps
27.794
peak memory bytes
69884854272
peak prompt tps
343.949
point count
224
ttft32k cached prompt tokens
28,668
ttft32k context tokens
32,768
ttft32k seconds
95.27

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25628.833UnknownUnknown59.994Unknown1
132,83227.794UnknownUnknown65.085Unknown1
1262,08019.814UnknownUnknown65.085Unknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:01a011ee-e409-783c-a752-119b6ceea579
repository
local.ai