Speed sweep

qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep

qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep

Record

Recipe
qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1
Measured
2026-07-04T14:59:57.814Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
132,76864.9245.4observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep
measured at
2026-07-04T14:59:57.814Z
recipe id
qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
llama.cpp qwen36 Spark build b831dc2fd962076c74d4b5837000459902cba484ee15f81d038d7d3ce1bcda62
latest point at
2026-07-04T14:59:57.814Z
max context tokens
32,768
peak generation tps
245.4
peak prompt tps
64.9
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
132,768245.4245.48,192Unknown64.91
Provenance & metadata (1)