Recipe
qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1
qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1MiaAI-Lab Qwen3.6-35B-A3B NVFP4 single-Spark vLLM profile with FlashInfer B12X, MTP-2, FP8 KV, and CUDA graphs
Record
- Status
- candidate
- Source
- mialabs
- Engine
- vllm
- Engine version
- 0.26+gb10
- Graph
- full-and-piecewise
- Accelerators
- 1
- Tensor parallel
- 1
- Max concurrency
- 8
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- yes
Hugging Face model card
Identity
https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4- Repository
- unsloth/Qwen3.6-35B-A3B-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
vllm
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/miaai-lab/mia-vllm-gb10-linear-b12x:latest@sha256:19627342e1da2607f4db50745dca30e57d7dd0ebff06062f03fd69b43a252931- Digest
sha256:19627342e1da2607f4db50745dca30e57d7dd0ebff06062f03fd69b43a252931- Port
- 8888
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
serveunsloth/Qwen3.6-35B-A3B-NVFP4--revision739af1e7aac320af1682ed1e0cce369af4c5265d--host0.0.0.0--port8888--tensor-parallel-size1--trust-remote-code--moe-backendauto--gpu-memory-utilization0.80--linear-backendflashinfer_b12x--attention-backendflashinfer--max-model-len262144--max-num-seqs24--max-num-batched-tokens32768--enable-chunked-prefill--async-scheduling--kv-cache-dtypefp8--limit-mm-per-prompt{"image":4}--allowed-media-domains*--speculative-config{"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"}--reasoning-parserqwen3--tool-call-parserqwen3_coder--enable-auto-tool-choice
| Flag | Value |
|---|---|
--revision | 739af1e7aac320af1682ed1e0cce369af4c5265d |
--host | 0.0.0.0 |
--port | 8888 |
--tensor-parallel-size | 1 |
--moe-backend | auto |
--gpu-memory-utilization | 0.80 |
--linear-backend | flashinfer_b12x |
--attention-backend | flashinfer |
--max-model-len | 262144 |
--max-num-seqs | 24 |
--max-num-batched-tokens | 32768 |
--kv-cache-dtype | fp8 |
--limit-mm-per-prompt | {"image":4} |
--allowed-media-domains | * |
--speculative-config | {"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"} |
--reasoning-parser | qwen3 |
--tool-call-parser | qwen3_coder |
Environment
| Variable | Value |
|---|---|
CUTE_DSL_ARCH | sm_121a |
HF_HOME | /root/.cache/huggingface |
Mounts
| Source | Target |
|---|---|
~/.cache/huggingface | /root/.cache/huggingface |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | — | 95.1 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
| 2 | — | — | 67.3 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
| 3 | — | — | 50.3 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
| 4 | — | — | 49.7 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
| 6 | — | — | 40.5 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
| 8 | — | — | 41.1 | — | historical | qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- dgx-spark-gb10-128gb
- id
- qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1
- model instance id
- unsloth-qwen3-6-35b-a3b-nvfp4--nvfp4
- recipe source
- mialabs
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- Yes
engine
- graph mode
- full-and-piecewise
- name
- vllm
- version
- 0.26+gb10
serving
- max concurrency
- 8
- tensor parallel
- 1
Provenance & metadata (3)
facts
- serving.kv cache tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max context tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |