Recipe

qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp4

qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp4

Local AI PostgreSQL candidate reconstructed from three GDPVal runs, all marked completed with infrastructure failures. The exact Qwen launch uses TP2, MTP4, 262K configured context, and 50 running sequences; model and image are pinned, but graph-mode, clean completion evidence, and speed sweeps remain missing, so CLI launch stays blocked.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
Graph
unknown
Accelerators
2
Tensor parallel
2
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4
Repository
nvidia/Qwen3.6-35B-A3B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

vllm

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/davidmcc73/vllm-openai@sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2
Digest
sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2
Port
8000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model
  2. nvidia/Qwen3.6-35B-A3B-NVFP4
  3. --served-model-name
  4. nvidia/Qwen3.6-35B-A3B-NVFP4
  5. --host
  6. 0.0.0.0
  7. --port
  8. 8000
  9. --tensor-parallel-size
  10. 2
  11. --pipeline-parallel-size
  12. 1
  13. --trust-remote-code
  14. --enable-chunked-prefill
  15. --enable-prefix-caching
  16. --enable-auto-tool-choice
  17. --enable-prompt-tokens-details
  18. --enable-force-include-usage
  19. --enable-request-id-headers
  20. --enable-log-requests
  21. --max-num-seqs
  22. 50
  23. --gpu-memory-utilization
  24. 0.85
  25. --block-size
  26. 1024
  27. --async-scheduling
  28. --enable-auto-tool-choice
  29. --language-model-only
  30. --attention-backend
  31. flashinfer
  32. --quantization
  33. modelopt
  34. --moe-backend
  35. marlin
  36. --max-model-len
  37. 262144
  38. --tool-call-parser
  39. qwen3_xml
  40. --reasoning-parser
  41. qwen3
  42. --kv-cache-dtype
  43. fp8
  44. --mamba-cache-mode
  45. align
  46. --speculative-config
  47. {"method":"mtp","num_speculative_tokens":4,"moe_backend":"triton","rejection_sample_method":"standard"}
FlagValue
--modelnvidia/Qwen3.6-35B-A3B-NVFP4
--served-model-namenvidia/Qwen3.6-35B-A3B-NVFP4
--host0.0.0.0
--port8000
--tensor-parallel-size2
--pipeline-parallel-size1
--max-num-seqs50
--gpu-memory-utilization0.85
--block-size1024
--attention-backendflashinfer
--quantizationmodelopt
--moe-backendmarlin
--max-model-len262144
--tool-call-parserqwen3_xml
--reasoning-parserqwen3
--kv-cache-dtypefp8
--mamba-cache-modealign
--speculative-config{"method":"mtp","num_speculative_tokens":4,"moe_backend":"triton","rejection_sample_method":"standard"}

Environment

VariableValue
FLASHINFER_DISABLE_VERSION_CHECK1
HF_HOME/root/.cache/huggingface
VLLM_FP8_MOE_BACKENDflashinfer_cutlass

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
rtx-pro-6000-blackwell-96gb
id
qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp4
model instance id
nvidia-qwen3-6-35b-a3b-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
unknown
name
vllm
version
0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96

serving

tensor parallel
2
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max concurrency · reason not-observed
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
speed sweep ids · provenance · captured at
2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry