Speed sweep
qwen36-35b-a3b-layered-fp8-arcb70-sglang-tp2-sweep
qwen36-35b-a3b-layered-fp8-arcb70-sglang-tp2-sweepRecord
- Measured
- 2026-08-25
- Accepted
- 2026-08-25
- Points
- 7
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | 128 | 58.4 | 115 | 2,190.2 | accepted |
| 1 | 512 | 58.9 | 116.2 | 8,687.3 | accepted |
| 1 | 2,048 | 62 | 113.9 | 33,060.1 | accepted |
| 1 | 8,192 | 62.6 | 116.4 | 130,857.6 | accepted |
| 2 | 512 | — | 85.7 | 15,724.6 | accepted |
| 4 | 512 | — | 67.8 | 33,694.9 | accepted |
| 8 | 512 | — | 41.8 | 66,861 | accepted |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-08-25
- id
- qwen36-35b-a3b-layered-fp8-arcb70-sglang-tp2-sweep
- measured at
- 2026-08-25
- recipe id
- qwen36-35b-a3b-layered-fp8-arcb70-sglang-tp2
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 8
- latest point at
- 2026-08-25
- max context tokens
- 8,192
- peak generation tps
- 116.35
- peak prompt tps
- 62.6
- point count
- 7
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| 1 | 128 | 114.95 | 114.95 | 128 | Unknown | 58.44 | 1 |
| 1 | 512 | 116.17 | 116.17 | 128 | Unknown | 58.94 | 1 |
| 1 | 2,048 | 113.92 | 113.92 | 128 | Unknown | 61.95 | 1 |
| 1 | 8,192 | 116.35 | 116.35 | 128 | Unknown | 62.6 | 1 |
| 2 | 512 | Unknown | 85.7 | 128 | Unknown | Unknown | 2 |
| 4 | 512 | Unknown | 67.8 | 128 | Unknown | Unknown | 4 |
| 8 | 512 | Unknown | 41.75 | 128 | Unknown | Unknown | 8 |
Provenance & metadata (1)
source
- commit
- 15166aed68bc25c411e16ddd81621242ba3d6298
- paths
- README.md, sglang/xpu/run_qwen3_6_service.sh, sglang/xpu/start_qwen3_6_service.sh, benchmarks/bench.py, benchmarks/results/20260825-195430-b70-sglang.json
- repository
- github.com/0xSero/qwen36-b70 ↗