Speed sweep

pg-019f3295-e23f-7045-a473-4ea475cd98a6-sweep

pg-019f3295-e23f-7045-a473-4ea475cd98a6-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-super-120b-a12b-4bit-4bit-apple-m3-ultra-96gb-80c-mlx-5e2708610b
Measured
2026-07-05T14:01:51.419Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25650.9observed
132,83247.6observed
1258,11132observed
132,76852,042observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f3295-e23f-7045-a473-4ea475cd98a6-sweep
measured at
2026-07-05T14:01:51.419Z
recipe id
pg-mlx-community-nvidia-nemotron-3-super-120b-a12b-4bit-4bit-apple-m3-ultra-96gb-80c-mlx-5e2708610b
schema version
local-ai-registry/v1

metrics

base memory bytes
69242544356
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
47.648
decode8k context tokens
8,256
decode8k tps
50.878
decode max context tokens
258,111
decode max context tps
31.957
decode mode
non-mtp
inference engine version
unknown
latest point at
2026-07-05T14:01:51.419Z
max context tokens
258,111
max prompt tokens
258,047
memory8k bytes
77217455288
memory8k context tokens
8,256
memory max context bytes
81244019960
memory max context tokens
258,111
peak generation tps
49.935
peak memory bytes
81244019960
peak prompt tps
38.592
point count
223
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
52.042

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25650.87850.878Unknown71.914Unknown1
132,83247.64847.648Unknown75.664Unknown1
1258,11131.95731.957Unknown75.664Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f3295-e23f-7045-a473-4ea475cd98a6
repository
local.ai