Speed sweep
qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-u760xh0t-sweep
qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-u760xh0t-sweepRecord
- Recipe
- qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-u760xh0t
- Measured
- 2026-07-06T14:11:52.053Z
- Points
- 1
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| — | 2,048 | — | 68.2 | 479.1 | observed |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- id
- qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-u760xh0t-sweep
- measured at
- 2026-07-06T14:11:52.053Z
- recipe id
- qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-u760xh0t
- schema version
- local-ai-registry/v1
metrics
- inference engine version
- 0.20.2rc1.dev13+g9557d9108.d20260620 local XPU patch stack
- latest point at
- 2026-07-06T14:11:52.053Z
- max context tokens
- 2,048
- peak generation tps
- 68.236
- point count
- 1
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| Unknown | 2,048 | 68.236 | Unknown | 512 | Unknown | Unknown | 1 |
Provenance & metadata (1)
source
- kind
- leaderboard
- repository
- www.localmaxxing.com ↗