Speed sweep

pg-019f2c34-e6d5-7593-9dc2-8a65bd000145-sweep

pg-019f2c34-e6d5-7593-9dc2-8a65bd000145-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m5-pro-64gb-mlx-f00534732e
Measured
2026-07-04T08:18:12.304Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25667.6observed
132,83260.7observed
1258,11132observed
132,76820,514.2observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f2c34-e6d5-7593-9dc2-8a65bd000145-sweep
measured at
2026-07-04T08:18:12.304Z
recipe id
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-nvfp4-nvfp4-apple-m5-pro-64gb-mlx-f00534732e
schema version
local-ai-registry/v1

metrics

base memory bytes
19923071196
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
60.651
decode8k context tokens
8,256
decode8k tps
67.585
decode max context tokens
258,111
decode max context tps
31.985
decode mode
non-mtp
inference engine version
unknown
latest point at
2026-07-04T08:18:12.304Z
max context tokens
258,111
max prompt tokens
258,047
memory8k bytes
22882245754
memory8k context tokens
8,256
memory max context bytes
26108318854
memory max context tokens
258,111
peak generation tps
60.651
peak memory bytes
26108318854
peak prompt tps
1,597.336
point count
223
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
20.514

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25667.58567.585Unknown21.311Unknown1
132,83260.65160.651Unknown24.315Unknown1
1258,11131.98531.985Unknown24.315Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f2c34-e6d5-7593-9dc2-8a65bd000145
repository
local.ai