Recipe

qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1

qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1

vllm recipe for Qwen3.6-27B-int4-AutoRound on Intel Arc Pro B70 with tensor parallelism 1.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
0.20.2rc1.dev13 local XPU patch stack
Accelerators
2
Tensor parallel
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/webhie/Qwen3.6-27B-int4-AutoRound
Repository
webhie/Qwen3.6-27B-int4-AutoRound
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmrgn3szj005dmj01u8tel6yd

No tokenized launch fields. This is measured or documented compatibility, not a Docker launch.

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
2,04887732.1historicalqwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
intel-arc-pro-b70-32gb
id
qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1
model instance id
webhie-qwen3-6-27b-int4-autoround--int4-autoround-w4a16-int8-head-int4-draft
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-6-27b-int4-autoround-w4a16-int8-head-int4-draft-intel-arc-pro-b70-32gb-vllm-tp1-sweep
status
candidate

capabilities

engine

name
vllm
version
0.20.2rc1.dev13 local XPU patch stack

serving

tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-27T06:04:13.773Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-27T06:04:13.773Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-27T06:04:13.773Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-27T06:04:13.773Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
description · provenance · captured at
2026-08-31T22:39:04Z
description · reason derived-from-recipe-identifiers
description · state known
engine.graph mode · provenance · captured at
2026-08-27T06:04:13.773Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
metadata.localmaxxing.engine flags.tensorParallel · provenance · captured at
2026-08-27T06:04:13.773Z
metadata.localmaxxing.engine flags.tensorParallel · reason not-observed
metadata.localmaxxing.engine flags.tensorParallel · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-09-01T10:20:39Z
serving.max context tokens · reason context-limit-not-evidenced
serving.max context tokens · state unknown

metadata

localmaxxing · base lineage
known
localmaxxing · engine flags · attentionBackend
vLLM XPU / Level Zero
localmaxxing · engine flags · flashAttn
No
localmaxxing · engine flags · gpuLayers
99
localmaxxing · engine flags · kvCacheDtype
f16
localmaxxing · engine flags · mtpEnabled
No
localmaxxing · engine flags · specDecoding
Yes
localmaxxing · hardware label
2x Intel Arc Pro B70 32GB
localmaxxing · notes
Strict fresh-response Qwen3.6 27B INT4 AutoRound TP2 record on two Intel Arc Pro B70 GPUs. Capturing all 48 GDN cores inside surrounding target PIECEWISE segments reduces target graph pieces from 129 to 33. Fixed 12-prompt realistic cold suite, cached_tokens=0 throughout, no cache/history/response reuse, target-verified MTP3, exact canaries + repeat128 + baseline parity + 1K needle all passed. Conservative full-quality result 87.029 tok/s; independent isolated high 87.816; swapped four-GPU crossover favored the candidate in both assignments.
localmaxxing · revision unpinned
No
localmaxxing · run id
cmrgn3szj005dmj01u8tel6yd

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipewww.localmaxxing.com/en/runs/cmrgn3szj005dmj01u8tel6yd