Speed sweep

pg-019f3295-e56f-7bbc-80e1-a80cae3f68f2-sweep

pg-019f3295-e56f-7bbc-80e1-a80cae3f68f2-sweep

Record

Recipe
pg-mlx-community-gpt-oss-120b-mxfp4-q8-mxfp4-q8-apple-m3-ultra-96gb-80c-mlx-ed712f6469
Measured
2026-07-05T14:01:52.235Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25566.8observed
132,83153.8observed
1127,03930.5observed
132,76724,831.8observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-019f3295-e56f-7bbc-80e1-a80cae3f68f2-sweep
measured at
2026-07-05T14:01:52.235Z
recipe id
pg-mlx-community-gpt-oss-120b-mxfp4-q8-mxfp4-q8-apple-m3-ultra-96gb-80c-mlx-ed712f6469
schema version
local-ai-registry/v1

metrics

base memory bytes
63553160102
base memory context tokens
191
concurrency
1
decode32k context tokens
32,831
decode32k tps
53.818
decode8k context tokens
8,255
decode8k tps
66.829
decode max context tokens
127,039
decode max context tps
30.475
decode mode
non-mtp
inference engine version
0.31.3
latest point at
2026-07-05T14:01:52.235Z
max context tokens
127,039
max prompt tokens
126,975
memory8k bytes
64824624762
memory8k context tokens
8,255
memory max context bytes
73115584934
memory max context tokens
127,039
peak generation tps
53.818
peak memory bytes
73115584934
peak prompt tps
1,319.556
point count
16
ttft32k cached prompt tokens
0
ttft32k context tokens
32,767
ttft32k seconds
24.832

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25566.82966.829Unknown60.373Unknown1
132,83153.81853.818Unknown68.094Unknown1
1127,03930.47530.475Unknown68.094Unknown1
132,767UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:019f3295-e56f-7bbc-80e1-a80cae3f68f2
repository
local.ai