Speed sweep

qwen3-8-27b-gptq-int4-intel-arc-pro-b70-32gb-vllm-tp2-sweep

qwen3-8-27b-gptq-int4-intel-arc-pro-b70-32gb-vllm-tp2-sweep

Record

Recipe
qwen3-8-27b-gptq-int4-intel-arc-pro-b70-32gb-vllm-tp2
Measured
2026-08-21T10:14:21.225Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
2,048710.3105.6891.2observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
qwen3-8-27b-gptq-int4-intel-arc-pro-b70-32gb-vllm-tp2-sweep
measured at
2026-08-21T10:14:21.225Z
recipe id
qwen3-8-27b-gptq-int4-intel-arc-pro-b70-32gb-vllm-tp2
schema version
local-ai-registry/v1

metrics

inference engine version
vLLM 0.26.1rc1.dev771+g8e6d8e4f6 XPU (editable tree, custom-patched: host-staged TP2 collectives + load-time INT8 W8A8 lm_head via oneDNN kernels)
latest point at
2026-08-21T10:14:21.225Z
max context tokens
2,048
peak generation tps
105.6
peak prompt tps
710.3
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
Unknown2,048105.6Unknown102Unknown710.31
Provenance & metadata (1)