Speed sweep
qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep
qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweepRecord
- Measured
- 2026-07-04T14:59:57.814Z
- Points
- 1
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 32,768 | 64.9 | 245.4 | — | observed |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- id
- qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep
- measured at
- 2026-07-04T14:59:57.814Z
- recipe id
- qwen3-5-35b-a3b-base-gguf-nvfp4-qwen35moe-dgx-spark-gb10-128gb-llama-cpp-tp1
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 1
- inference engine version
- llama.cpp qwen36 Spark build b831dc2fd962076c74d4b5837000459902cba484ee15f81d038d7d3ce1bcda62
- latest point at
- 2026-07-04T14:59:57.814Z
- max context tokens
- 32,768
- peak generation tps
- 245.4
- peak prompt tps
- 64.9
- point count
- 1
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| 1 | 32,768 | 245.4 | 245.4 | 8,192 | Unknown | 64.9 | 1 |
Provenance & metadata (1)
source
- kind
- leaderboard
- repository
- www.localmaxxing.com ↗