Speed sweep

pg-019f90a8-e82c-7565-82da-b8a7f0339d46-sweep

pg-019f90a8-e82c-7565-82da-b8a7f0339d46-sweep

Record

Recipe
pg-unsloth-nvidia-nemotron-3-super-120b-a12b-gguf-ud-q3-k-m-nv-a8bc4e5fe708-apple-m3-ultra-96gb-80c-llama-cpp-35cfe20ca2
Measured
2026-07-23T20:26:56.234Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25637.5observed
132,83236.5observed
1262,08028.2observed
132,76856,572.5observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f90a8-e82c-7565-82da-b8a7f0339d46-sweep
measured at
2026-07-23T20:26:56.234Z
recipe id
pg-unsloth-nvidia-nemotron-3-super-120b-a12b-gguf-ud-q3-k-m-nv-a8bc4e5fe708-apple-m3-ultra-96gb-80c-llama-cpp-35cfe20ca2
schema version
local-ai-registry/v1

metrics

base memory bytes
63388762112
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
36.481
decode8k context tokens
8,256
decode8k tps
37.527
decode max context tokens
262,080
decode max context tps
28.152
decode mode
non-mtp
inference engine version
unknown
latest point at
2026-07-23T20:26:56.234Z
max context tokens
262,080
max prompt tokens
262,016
memory8k bytes
64664174592
memory8k context tokens
8,256
memory max context bytes
70170624000
memory max context tokens
262,080
peak generation tps
36.481
peak memory bytes
70170624000
peak prompt tps
579.221
point count
224
ttft32k cached prompt tokens
28,668
ttft32k context tokens
32,768
ttft32k seconds
56.573

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25637.52737.527Unknown60.223Unknown1
132,83236.48136.481Unknown65.351Unknown1
1262,08028.15228.152Unknown65.351Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f90a8-e82c-7565-82da-b8a7f0339d46
repository
local.ai