Speed sweep

pg-019f45bd-5d78-7776-a150-03693e10e2ac-sweep

pg-019f45bd-5d78-7776-a150-03693e10e2ac-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m4-max-128gb-mlx-aeab2af3b9
Measured
2026-07-09T07:17:45.658Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25698.2observed
132,83287.1observed
1262,07943.8observed
132,76828,218observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f45bd-5d78-7776-a150-03693e10e2ac-sweep
measured at
2026-07-09T07:17:45.658Z
recipe id
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m4-max-128gb-mlx-aeab2af3b9
schema version
local-ai-registry/v1

metrics

base memory bytes
38701957412
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
87.076
decode8k context tokens
8,256
decode8k tps
98.244
decode max context tokens
262,079
decode max context tps
43.812
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-07-09T07:17:45.658Z
max context tokens
262,079
max prompt tokens
262,015
memory8k bytes
38701957412
memory8k context tokens
8,256
memory max context bytes
38701957412
memory max context tokens
262,079
peak generation tps
87.076
peak memory bytes
38701957412
peak prompt tps
1,161.243
point count
224
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
28.218

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25698.24498.244Unknown36.044Unknown1
132,83287.07687.076Unknown36.044Unknown1
1262,07943.81243.812Unknown36.044Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f45bd-5d78-7776-a150-03693e10e2ac
repository
local.ai