Speed sweep

pg-01a013a0-0851-710c-9573-e9cb90e0e737-sweep

pg-01a013a0-0851-710c-9573-e9cb90e0e737-sweep

Record

Recipe
pg-unsloth-deepseek-v4-flash-gguf-ud-q2-k-xl-deepseek-v4-flash-fef3949f4528-apple-m5-max-128gb-llama-cpp-b684067f9c
Measured
2026-08-18T06:47:30.121Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25613observed
132,83212.3observed
1237,6327.9observed
132,768162,980.1observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-01a013a0-0851-710c-9573-e9cb90e0e737-sweep
measured at
2026-08-18T06:47:30.121Z
recipe id
pg-unsloth-deepseek-v4-flash-gguf-ud-q2-k-xl-deepseek-v4-flash-fef3949f4528-apple-m5-max-128gb-llama-cpp-b684067f9c
schema version
local-ai-registry/v1

metrics

base memory bytes
99041525760
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
12.303
decode8k context tokens
8,256
decode8k tps
13.013
decode max context tokens
237,632
decode max context tps
7.912
decode mode
non-mtp
inference engine version
10017 (7dc1bfead)
latest point at
2026-08-18T06:47:30.121Z
max context tokens
237,632
max prompt tokens
237,568
memory8k bytes
99099459584
memory8k context tokens
8,256
memory max context bytes
99715842048
memory max context tokens
237,632
peak generation tps
12.303
peak memory bytes
99715842048
peak prompt tps
201.055
point count
63
ttft32k cached prompt tokens
8,188
ttft32k context tokens
32,768
ttft32k seconds
162.98

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25613.01313.013Unknown92.294Unknown1
132,83212.30312.303Unknown92.868Unknown1
1237,6327.9127.912Unknown92.868Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:01a013a0-0851-710c-9573-e9cb90e0e737
repository
local.ai