Recipe

qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1

qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1

MiaAI-Lab Qwen3.6-35B-A3B NVFP4 single-Spark vLLM profile with FlashInfer B12X, MTP-2, FP8 KV, and CUDA graphs

Record

Status
candidate
Source
mialabs
Engine
vllm
Engine version
0.26+gb10
Graph
full-and-piecewise
Accelerators
1
Tensor parallel
1
Max concurrency
8
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/unsloth/Qwen3.6-35B-A3B-NVFP4
Repository
unsloth/Qwen3.6-35B-A3B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

vllm

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/miaai-lab/mia-vllm-gb10-linear-b12x:latest@sha256:19627342e1da2607f4db50745dca30e57d7dd0ebff06062f03fd69b43a252931
Digest
sha256:19627342e1da2607f4db50745dca30e57d7dd0ebff06062f03fd69b43a252931
Port
8888

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. serve
  2. unsloth/Qwen3.6-35B-A3B-NVFP4
  3. --revision
  4. 739af1e7aac320af1682ed1e0cce369af4c5265d
  5. --host
  6. 0.0.0.0
  7. --port
  8. 8888
  9. --tensor-parallel-size
  10. 1
  11. --trust-remote-code
  12. --moe-backend
  13. auto
  14. --gpu-memory-utilization
  15. 0.80
  16. --linear-backend
  17. flashinfer_b12x
  18. --attention-backend
  19. flashinfer
  20. --max-model-len
  21. 262144
  22. --max-num-seqs
  23. 24
  24. --max-num-batched-tokens
  25. 32768
  26. --enable-chunked-prefill
  27. --async-scheduling
  28. --kv-cache-dtype
  29. fp8
  30. --limit-mm-per-prompt
  31. {"image":4}
  32. --allowed-media-domains
  33. *
  34. --speculative-config
  35. {"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"}
  36. --reasoning-parser
  37. qwen3
  38. --tool-call-parser
  39. qwen3_coder
  40. --enable-auto-tool-choice
FlagValue
--revision739af1e7aac320af1682ed1e0cce369af4c5265d
--host0.0.0.0
--port8888
--tensor-parallel-size1
--moe-backendauto
--gpu-memory-utilization0.80
--linear-backendflashinfer_b12x
--attention-backendflashinfer
--max-model-len262144
--max-num-seqs24
--max-num-batched-tokens32768
--kv-cache-dtypefp8
--limit-mm-per-prompt{"image":4}
--allowed-media-domains*
--speculative-config{"method":"mtp","num_speculative_tokens":2,"moe_backend":"triton"}
--reasoning-parserqwen3
--tool-call-parserqwen3_coder

Environment

VariableValue
CUTE_DSL_ARCHsm_121a
HF_HOME/root/.cache/huggingface

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
195.1historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
267.3historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
350.3historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
449.7historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
640.5historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
841.1historicalqwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1
model instance id
unsloth-qwen3-6-35b-a3b-nvfp4--nvfp4
recipe source
mialabs
schema version
local-ai-registry/v1
speed sweep ids
qwen36-35b-a3b-nvfp4-dgxspark-vllm-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
full-and-piecewise
name
vllm
version
0.26+gb10

serving

max concurrency
8
tensor parallel
1
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry