Speed sweep

pg-019f2c34-df8f-7aa7-acaf-103dcfd5b9a3-sweep

pg-019f2c34-df8f-7aa7-acaf-103dcfd5b9a3-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-mlx-8bit-8bit-apple-m5-max-128gb-mlx-e629aea6ea
Measured
2026-07-04T08:18:10.443Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,256100.5observed
132,83291.8observed
1258,11153.7observed
132,76811,057observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f2c34-df8f-7aa7-acaf-103dcfd5b9a3-sweep
measured at
2026-07-04T08:18:10.443Z
recipe id
pg-mlx-community-nvidia-nemotron-3-nano-30b-a3b-mlx-8bit-8bit-apple-m5-max-128gb-mlx-e629aea6ea
schema version
local-ai-registry/v1

metrics

base memory bytes
34159707128
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
91.768
decode8k context tokens
8,256
decode8k tps
100.451
decode max context tokens
258,111
decode max context tps
53.678
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-07-04T08:18:10.443Z
max context tokens
258,111
max prompt tokens
258,047
memory8k bytes
37118863226
memory8k context tokens
8,256
memory max context bytes
40344969094
memory max context tokens
258,111
peak generation tps
91.768
peak memory bytes
40344969094
peak prompt tps
2,963.561
point count
223
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
11.057

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,256100.451100.451Unknown34.57Unknown1
132,83291.76891.768Unknown37.574Unknown1
1258,11153.67853.678Unknown37.574Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f2c34-df8f-7aa7-acaf-103dcfd5b9a3
repository
local.ai