Speed sweep

pg-019f3295-e596-7117-bedc-b34b2bc69c6f-sweep

pg-019f3295-e596-7117-bedc-b34b2bc69c6f-sweep

Record

Recipe
pg-mlx-community-gpt-oss-20b-mxfp4-q8-mxfp4-q8-apple-m3-ultra-96gb-80c-mlx-3d193930e3
Measured
2026-07-05T14:01:52.273Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
12,111107observed
132,83179.4observed
1127,03945.9observed
132,76715,876observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f3295-e596-7117-bedc-b34b2bc69c6f-sweep
measured at
2026-07-05T14:01:52.273Z
recipe id
pg-mlx-community-gpt-oss-20b-mxfp4-q8-mxfp4-q8-apple-m3-ultra-96gb-80c-mlx-3d193930e3
schema version
local-ai-registry/v1

metrics

base memory bytes
12284006990
base memory context tokens
191
concurrency
1
decode32k context tokens
32,831
decode32k tps
79.432
decode8k context tokens
2,111
decode8k tps
106.981
decode max context tokens
127,039
decode max context tps
45.852
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-07-05T14:01:52.273Z
max context tokens
127,039
max prompt tokens
126,975
memory8k bytes
13005211082
memory8k context tokens
2,111
memory max context bytes
18783372482
memory max context tokens
127,039
peak generation tps
79.432
peak memory bytes
18783372482
peak prompt tps
2,063.938
point count
16
ttft32k cached prompt tokens
0
ttft32k context tokens
32,767
ttft32k seconds
15.876

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
12,111106.981106.981Unknown12.112Unknown1
132,83179.43279.432Unknown17.493Unknown1
1127,03945.85245.852Unknown17.493Unknown1
132,767UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f3295-e596-7117-bedc-b34b2bc69c6f
repository
local.ai