Speed sweep

pg-01a0303f-9dee-7a8c-b68b-6b91fe0a2a06-sweep

pg-01a0303f-9dee-7a8c-b68b-6b91fe0a2a06-sweep

Record

Recipe
pg-mlx-community-nvidia-nemotron-3-super-120b-a12b-4bit-4bit-apple-m3-ultra-96gb-60c-mlx-0b7c9789a9
Measured
2026-08-23T20:11:10.684Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25647.7observed
132,83246.1observed
1262,08022.9observed
132,76869,035.8observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-01a0303f-9dee-7a8c-b68b-6b91fe0a2a06-sweep
measured at
2026-08-23T20:11:10.684Z
recipe id
pg-mlx-community-nvidia-nemotron-3-super-120b-a12b-4bit-4bit-apple-m3-ultra-96gb-60c-mlx-0b7c9789a9
schema version
local-ai-registry/v1

metrics

base memory bytes
69249759887
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
46.107
decode8k context tokens
8,256
decode8k tps
47.658
decode max context tokens
262,080
decode max context tps
22.868
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-08-23T20:11:10.684Z
max context tokens
262,080
max prompt tokens
262,016
memory8k bytes
77217455288
memory8k context tokens
8,256
memory max context bytes
81378254148
memory max context tokens
262,080
peak generation tps
46.107
peak memory bytes
81378254148
peak prompt tps
474.653
point count
224
ttft32k cached prompt tokens
28,736
ttft32k context tokens
32,768
ttft32k seconds
69.036

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25647.65847.658Unknown71.914Unknown1
132,83246.10746.107Unknown75.789Unknown1
1262,08022.86822.868Unknown75.789Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:01a0303f-9dee-7a8c-b68b-6b91fe0a2a06
repository
local.ai