Recipe
qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2
qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- vllm
- Engine version
- 0.20.2rc1.dev2+gc51df4300.d20260523
- Graph
- piecewise
- Accelerators
- 2
- Tensor parallel
- 2
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8- Repository
- nameistoken/Qwen3.6-35B-A3B-Quark-W8A8-INT8
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
vllm
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
bashscripts/run-qwen36-ablation-candidate.shprefill-safe-int8-mixed-workspace-async-tp2-smoke
Environment
| Variable | Value |
|---|---|
ABLATION_RUN_QUALITY | 0 |
COLOR_REPEATS | 16 |
COMPILATION_CONFIG | {"cudagraph_mode":"PIECEWISE"} |
GPU_MEMORY_UTILIZATION | 0.90 |
JSON_REPEATS | 16 |
MAX_NUM_SEQS | 24 |
METRICS_REPEATS | 1 |
ONEAPI_DEVICE_SELECTOR | level_zero:0,1 |
SERVER_LAUNCHER | scripts/launch-qwen36-quark-int8-accepted.sh |
TP_SIZE | 2 |
VLLM_EXTRA_ARGS | --uvicorn-log-level warning |
VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY | 1 |
VLLM_XPU_ENABLE_XPU_GRAPH | 1 |
VLLM_XPU_FORCE_GRAPH_WITH_COMM | 1 |
VLLM_XPU_GDN_NATIVE_FALLBACK | prefill |
VLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK | 1 |
VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE | 1 |
VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK | 1 |
VLLM_XPU_INT8_MOE_MIXED_WORKSPACE | 1 |
XPU_GRAPH | 1 |
ZE_AFFINITY_MASK | 0,1 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| — | 32,768 | — | 85.9 | 272.8 | observed | qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 2
- hardware id
- intel-arc-pro-b70-32gb
- id
- qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2
- model instance id
- nameistoken-qwen3-6-35b-a3b-quark-w8a8-int8--quark-w8a8-int8
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen3-6-35b-a3b-quark-w8a8-int8-intel-arc-pro-b70-32gb-vllm-tp2-sweep
- status
- candidate
capabilities
engine
- graph mode
- piecewise
- name
- vllm
- version
- 0.20.2rc1.dev2+gc51df4300.d20260523
serving
- tensor parallel
- 2
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-09-01T01:32:29Z
engine.graph mode · reason explicit-cudagraph-mode-in-observed-command
engine.graph mode · state known
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max concurrency · provenance · captured at
- 2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown
- serving.max context tokens · provenance · captured at
- 2026-09-01T10:20:39Z
serving.max context tokens · reason context-limit-not-evidenced
serving.max context tokens · state unknown
metadata
- localmaxxing · backend
- xpu
- localmaxxing · batch size
- 1
- localmaxxing · hardware label
- Intel Arc Pro B70
- localmaxxing · notes
- Best safer TP2 reference for Qwen3.6 35B A3B Quark W8A8 INT8 on 2x Intel Arc Pro B70. Shape p512/o512 through OpenAI-compatible vLLM/XPU endpoint, temperature 0, corrected output throughput after first streamed chunk. Validation: JSON canary 16/16 pass and color/order canary 16/16 pass; quality suite was skipped, so this is a safe-smoke/reference submission rather than the stronger TP4 deep gate. Older TP2 raw metrics reached about 91 tok/s but were not promoted under the current gates. GitHub result packet: https://github.com/steveseguin/b70-optimization-lab/blob/main/results/qwen36-35b-quark-int8-b70/2x-b70-reference.md . Evidence: data/qwen36-ablation-prefill-safe-int8-mixed-workspace-async-tp2-smoke-summary-20260615tp2safe1.json and data/qwen36-ablation-prefill-safe-int8-mixed-workspace-async-tp2-smoke-p512o512-20260615tp2safe1.json.
- localmaxxing · observed command
- SERVER_LAUNCHER=scripts/launch-qwen36-quark-int8-accepted.sh TP_SIZE=2 ONEAPI_DEVICE_SELECTOR=level_zero:0,1 ZE_AFFINITY_MASK=0,1 XPU_GRAPH=1 VLLM_XPU_ENABLE_XPU_GRAPH=1 VLLM_XPU_FORCE_GRAPH_WITH_COMM=1 VLLM_XPU_GRAPH_NOOP_COMM_CAPTURE=1 COMPILATION_CONFIG='{"cudagraph_mode":"PIECEWISE"}' VLLM_XPU_GDN_NATIVE_FALLBACK=prefill VLLM_XPU_GDN_PREFILL_RECURRENT_FALLBACK=1 VLLM_XPU_DISABLE_PREFILL_CUDAGRAPH_REPLAY=1 VLLM_XPU_GREEDY_SAMPLE_TOPK_FALLBACK=1 VLLM_XPU_INT8_MOE_MIXED_WORKSPACE=1 GPU_MEMORY_UTILIZATION=0.90 MAX_NUM_SEQS=24 VLLM_EXTRA_ARGS='--uvicorn-log-level warning' METRICS_REPEATS=1 JSON_REPEATS=16 COLOR_REPEATS=16 ABLATION_RUN_QUALITY=0 bash scripts/run-qwen36-ablation-candidate.sh prefill-safe-int8-mixed-workspace-async-tp2-smoke
- localmaxxing · run id
- cmqq4mwgm00yiqo0133bj962q
- localmaxxing · tokenized · arguments
- bash, scripts/run-qwen36-ablation-candidate.sh, prefill-safe-int8-mixed-workspace-async-tp2-smoke
- localmaxxing · tokenized · environment · ABLATION RUN QUALITY
- 0
- localmaxxing · tokenized · environment · COLOR REPEATS
- 16
- localmaxxing · tokenized · environment · COMPILATION CONFIG
- {"cudagraph_mode":"PIECEWISE"}
- localmaxxing · tokenized · environment · GPU MEMORY UTILIZATION
- 0.90
- localmaxxing · tokenized · environment · JSON REPEATS
- 16
- localmaxxing · tokenized · environment · MAX NUM SEQS
- 24
- localmaxxing · tokenized · environment · METRICS REPEATS
- 1
- localmaxxing · tokenized · environment · ONEAPI DEVICE SELECTOR
- level_zero:0,1
- localmaxxing · tokenized · environment · SERVER LAUNCHER
- scripts/launch-qwen36-quark-int8-accepted.sh
- localmaxxing · tokenized · environment · TP SIZE
- 2
- localmaxxing · tokenized · environment · VLLM EXTRA ARGS
- --uvicorn-log-level warning
- localmaxxing · tokenized · environment · VLLM XPU DISABLE PREFILL CUDAGRAPH REPLAY
- 1
- localmaxxing · tokenized · environment · VLLM XPU ENABLE XPU GRAPH
- 1
- localmaxxing · tokenized · environment · VLLM XPU FORCE GRAPH WITH COMM
- 1
- localmaxxing · tokenized · environment · VLLM XPU GDN NATIVE FALLBACK
- prefill
- localmaxxing · tokenized · environment · VLLM XPU GDN PREFILL RECURRENT FALLBACK
- 1
- localmaxxing · tokenized · environment · VLLM XPU GRAPH NOOP COMM CAPTURE
- 1
- localmaxxing · tokenized · environment · VLLM XPU GREEDY SAMPLE TOPK FALLBACK
- 1
- localmaxxing · tokenized · environment · VLLM XPU INT8 MOE MIXED WORKSPACE
- 1
- localmaxxing · tokenized · environment · XPU GRAPH
- 1
- localmaxxing · tokenized · environment · ZE AFFINITY MASK
- 0,1
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmqq4mwgm00yiqo0133bj962q ↗ |