Recipe

qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2

qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
0.20.2rc1.dev2+gc51df4300.d20260523
Graph
piecewise
Accelerators
2
Tensor parallel
2
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8
Repository
nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmqq4mwgm00yiqo0133bj962q

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. bash
  2. scripts/run-qwen36-ablation-candidate.sh
  3. prefill-safe-int8-mixed-workspace-async-tp2-smoke

Environment

VariableValue
ABLATION_RUN_QUALITY0
COLOR_REPEATS16
COMPILATION_CONFIG{"cudagraph_mode":"PIECEWISE"}
GPU_MEMORY_UTILIZATION0.90
JSON_REPEATS16
MAX_NUM_SEQS24
METRICS_REPEATS1
ONEAPI_DEVICE_SELECTORlevel_zero:0,1
SERVER_LAUNCHERscripts/launch-qwen36-quark-int8-accepted.sh
TP_SIZE2
VLLM_EXTRA_ARGS--uvicorn-log-level warning
VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY1
VLLM_XPU_ENABLE_XPU_GRAPH1
VLLM_XPU_FORCE_GRAPH_WITH_COMM1
VLLM_XPU_GDN_NATIVE_FALLBACKprefill
VLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK1
VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE1
VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK1
VLLM_XPU_INT8_MOE_MIXED_WORKSPACE1
XPU_GRAPH1
ZE_AFFINITY_MASK0,1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
32,76885.9272.8observedqwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
intel-arc-pro-b70-32gb
id
qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2
model instance id
nameistoken-qwen3-6-35b-a3b-quark-w8a8-int8--quark-w8a8-int8
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2-sweep
status
candidate

capabilities

engine

graph mode
piecewise
name
vllm
version
0.20.2rc1.dev2+gc51df4300.d20260523

serving

tensor parallel
2
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-09-01T01:32:29Z
engine.graph mode · reason explicit-cudagraph-mode-in-observed-command
engine.graph mode · state known
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-09-01T10:20:39Z
serving.max context tokens · reason context-limit-not-evidenced
serving.max context tokens · state unknown

metadata

localmaxxing · backend
xpu
localmaxxing · batch size
1
localmaxxing · hardware label
Intel Arc Pro B70
localmaxxing · notes
Best safer TP2 reference for Qwen3.6 35B A3B Quark W8A8 INT8 on 2x Intel Arc Pro B70. Shape p512/o512 through OpenAI-compatible vLLM/XPU endpoint, temperature 0, corrected output throughput after first streamed chunk. Validation: JSON canary 16/16 pass and color/order canary 16/16 pass; quality suite was skipped, so this is a safe-smoke/reference submission rather than the stronger TP4 deep gate. Older TP2 raw metrics reached about 91 tok/s but were not promoted under the current gates. GitHub result packet: https://github.com/steveseguin/b70-optimization-lab/blob/main/results/qwen36-35b-quark-int8-b70/2x-b70-reference.md . Evidence: data/qwen36-ablation-prefill-safe-int8-mixed-workspace-async-tp2-smoke-summary-20260615tp2safe1.json and data/qwen36-ablation-prefill-safe-int8-mixed-workspace-async-tp2-smoke-p512o512-20260615tp2safe1.json.
localmaxxing · observed command
SERVER_LAUNCHER=scripts/launch-qwen36-quark-int8-accepted.sh TP_SIZE=2 ONEAPI_DEVICE_SELECTOR=level_zero:0,1 ZE_AFFINITY_MASK=0,1 XPU_GRAPH=1 VLLM_XPU_ENABLE_XPU_GRAPH=1 VLLM_XPU_FORCE_GRAPH_WITH_COMM=1 VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE=1 COMPILATION_CONFIG='{"cudagraph_mode":"PIECEWISE"}' VLLM_XPU_GDN_NATIVE_FALLBACK=prefill VLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK=1 VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY=1 VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK=1 VLLM_XPU_INT8_MOE_MIXED_WORKSPACE=1 GPU_MEMORY_UTILIZATION=0.90 MAX_NUM_SEQS=24 VLLM_EXTRA_ARGS='--uvicorn-log-level warning' METRICS_REPEATS=1 JSON_REPEATS=16 COLOR_REPEATS=16 ABLATION_RUN_QUALITY=0 bash scripts/run-qwen36-ablation-candidate.sh prefill-safe-int8-mixed-workspace-async-tp2-smoke
localmaxxing · run id
cmqq4mwgm00yiqo0133bj962q
localmaxxing · tokenized · arguments
bash, scripts/run-qwen36-ablation-candidate.sh, prefill-safe-int8-mixed-workspace-async-tp2-smoke
localmaxxing · tokenized · environment · ABLATION RUN QUALITY
0
localmaxxing · tokenized · environment · COLOR REPEATS
16
localmaxxing · tokenized · environment · COMPILATION CONFIG
{"cudagraph_mode":"PIECEWISE"}
localmaxxing · tokenized · environment · GPU MEMORY UTILIZATION
0.90
localmaxxing · tokenized · environment · JSON REPEATS
16
localmaxxing · tokenized · environment · MAX NUM SEQS
24
localmaxxing · tokenized · environment · METRICS REPEATS
1
localmaxxing · tokenized · environment · ONEAPI DEVICE SELECTOR
level_zero:0,1
localmaxxing · tokenized · environment · SERVER LAUNCHER
scripts/launch-qwen36-quark-int8-accepted.sh
localmaxxing · tokenized · environment · TP SIZE
2
localmaxxing · tokenized · environment · VLLM EXTRA ARGS
--uvicorn-log-level warning
localmaxxing · tokenized · environment · VLLM XPU DISABLE PREFILL CUDAGRAPH REPLAY
1
localmaxxing · tokenized · environment · VLLM XPU ENABLE XPU GRAPH
1
localmaxxing · tokenized · environment · VLLM XPU FORCE GRAPH WITH COMM
1
localmaxxing · tokenized · environment · VLLM XPU GDN NATIVE FALLBACK
prefill
localmaxxing · tokenized · environment · VLLM XPU GDN PREFILL RECURRENT FALLBACK
1
localmaxxing · tokenized · environment · VLLM XPU GRAPH NOOP COMM CAPTURE
1
localmaxxing · tokenized · environment · VLLM XPU GREEDY SAMPLE TOPK FALLBACK
1
localmaxxing · tokenized · environment · VLLM XPU INT8 MOE MIXED WORKSPACE
1
localmaxxing · tokenized · environment · XPU GRAPH
1
localmaxxing · tokenized · environment · ZE AFFINITY MASK
0,1
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmqq4mwgm00yiqo0133bj962q