Speed sweep

pg-01a026ab-730a-75e5-9a74-f6be86e2acda-sweep

pg-01a026ab-730a-75e5-9a74-f6be86e2acda-sweep

Record

Recipe
pg-unsloth-deepseek-v4-flash-gguf-ud-q2-k-xl-deepseek-v4-flash-fef3949f4528-apple-m5-max-128gb-llama-cpp-b684067f9c
Measured
2026-08-21T23:32:45.435Z
Points
4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
18,25632.2observed
132,83229.1observed
1262,08020.5observed
132,768141,932observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
pg-01a026ab-730a-75e5-9a74-f6be86e2acda-sweep
measured at
2026-08-21T23:32:45.435Z
recipe id
pg-unsloth-deepseek-v4-flash-gguf-ud-q2-k-xl-deepseek-v4-flash-fef3949f4528-apple-m5-max-128gb-llama-cpp-b684067f9c
schema version
local-ai-registry/v1

metrics

base memory bytes
99013951488
base memory context tokens
192
concurrency
1
decode32k context tokens
32,832
decode32k tps
29.106
decode8k context tokens
8,256
decode8k tps
32.154
decode max context tokens
262,080
decode max context tps
20.521
decode mode
non-mtp
inference engine version
10276 (6ea215d)
latest point at
2026-08-21T23:32:45.435Z
max context tokens
262,080
max prompt tokens
262,016
memory8k bytes
99067723776
memory8k context tokens
8,256
memory max context bytes
99907010560
memory max context tokens
262,080
peak generation tps
29.338
peak memory bytes
99907010560
peak prompt tps
245.442
point count
224
ttft32k cached prompt tokens
28,735
ttft32k context tokens
32,768
ttft32k seconds
141.932

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
18,25632.15432.154Unknown92.264Unknown1
132,83229.10629.106Unknown93.046Unknown1
1262,08020.52120.521Unknown93.046Unknown1
132,768UnknownUnknownUnknownUnknownUnknown1
Provenance & metadata (1)

source

paths
publication:pg-20260827T060320709Z, run:01a026ab-730a-75e5-9a74-f6be86e2acda
repository
local.ai