Speed sweep

pg-019f2c34-cff5-7d88-893d-5a7ed6218d18-sweep

pg-019f2c34-cff5-7d88-893d-5a7ed6218d18-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m3-ultra-96gb-60c-mlx-7e0d994571
Measured
2026-07-04T08:18:06.448Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25698.3observed
132,83289.3observed
1258,11148observed
132,76822,434.8observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f2c34-cff5-7d88-893d-5a7ed6218d18-sweep
measured at
2026-07-04T08:18:06.448Z
recipe id
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m3-ultra-96gb-60c-mlx-7e0d994571
schema version
local-ai-registry/v1

metrics

base memory bytes
19924908528
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
89.315
decode8k context tokens
8,256
decode8k tps
98.282
decode max context tokens
258,111
decode max context tps
48
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-07-04T08:18:06.448Z
max context tokens
258,111
max prompt tokens
258,047
memory8k bytes
22882262138
memory8k context tokens
8,256
memory max context bytes
26110334086
memory max context tokens
258,111
peak generation tps
89.315
peak memory bytes
26110334086
peak prompt tps
1,460.587
point count
223
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
22.435

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25698.28298.282Unknown21.311Unknown1
132,83289.31589.315Unknown24.317Unknown1
1258,1114848Unknown24.317Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f2c34-cff5-7d88-893d-5a7ed6218d18
repository
local.ai