Speed sweep

qwen3-5-35b-a3b-iq1-m-rtx-5070-12gb-llama-cpp-tp1-sweep

qwen3-5-35b-a3b-iq1-m-rtx-5070-12gb-llama-cpp-tp1-sweep

Record

Recipe
qwen3-5-35b-a3b-iq1-m-rtx-5070-12gb-llama-cpp-tp1
Measured
2026-08-26T03:09:59.219Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
12,0481,730.9182.3295.8observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
qwen3-5-35b-a3b-iq1-m-rtx-5070-12gb-llama-cpp-tp1-sweep
measured at
2026-08-26T03:09:59.219Z
recipe id
qwen3-5-35b-a3b-iq1-m-rtx-5070-12gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
llama.cpp f280b26983ad0fdb705a0d9ebf0503e76f2899b0 (campaign build sm_120, CUDA 13.3)
latest point at
2026-08-26T03:09:59.219Z
max context tokens
2,048
peak generation tps
182.31
peak prompt tps
1,730.93
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
12,048182.31182.31128Unknown1,730.931
Provenance & metadata (1)