Recipe
qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp3
qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp3Local AI PostgreSQL candidate reconstructed from completed GAIA, GDPVal, and TAU2 evaluations. The exact Qwen launch uses TP2, MTP3, 262K configured context, and 51 running sequences; model and image are pinned, but graph-mode, an on-hardware completion artifact, and speed sweeps remain missing, so CLI launch stays blocked.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- 0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
- Graph
- unknown
- Accelerators
- 2
- Tensor parallel
- 2
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/nvidia/Qwen3.6-35B-A3B-NVFP4- Repository
- nvidia/Qwen3.6-35B-A3B-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
vllm
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/davidmcc73/vllm-openai@sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2- Digest
sha256:f03040c06dd43c0b48d0b471ae67edcc0c8fe8e63d4f762489d7be0015f527b2- Port
- 8000
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--modelnvidia/Qwen3.6-35B-A3B-NVFP4--served-model-namenvidia/Qwen3.6-35B-A3B-NVFP4--host0.0.0.0--port8000--tensor-parallel-size2--pipeline-parallel-size1--trust-remote-code--enable-chunked-prefill--enable-prefix-caching--enable-auto-tool-choice--enable-prompt-tokens-details--enable-force-include-usage--enable-request-id-headers--enable-log-requests--max-num-seqs51--gpu-memory-utilization0.85--block-size1024--async-scheduling--enable-auto-tool-choice--language-model-only--attention-backendflashinfer--quantizationmodelopt--moe-backendmarlin--max-model-len262144--tool-call-parserqwen3_xml--reasoning-parserqwen3--kv-cache-dtypefp8--mamba-cache-modealign--speculative-config{"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton","rejection_sample_method":"standard"}
| Flag | Value |
|---|---|
--model | nvidia/Qwen3.6-35B-A3B-NVFP4 |
--served-model-name | nvidia/Qwen3.6-35B-A3B-NVFP4 |
--host | 0.0.0.0 |
--port | 8000 |
--tensor-parallel-size | 2 |
--pipeline-parallel-size | 1 |
--max-num-seqs | 51 |
--gpu-memory-utilization | 0.85 |
--block-size | 1024 |
--attention-backend | flashinfer |
--quantization | modelopt |
--moe-backend | marlin |
--max-model-len | 262144 |
--tool-call-parser | qwen3_xml |
--reasoning-parser | qwen3 |
--kv-cache-dtype | fp8 |
--mamba-cache-mode | align |
--speculative-config | {"method":"mtp","num_speculative_tokens":3,"moe_backend":"triton","rejection_sample_method":"standard"} |
Environment
| Variable | Value |
|---|---|
FLASHINFER_DISABLE_VERSION_CHECK | 1 |
HF_HOME | /root/.cache/huggingface |
VLLM_FP8_MOE_BACKEND | flashinfer_cutlass |
Mounts
| Source | Target |
|---|---|
~/.cache/huggingface | /root/.cache/huggingface |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 2
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- qwen3.6-35b-a3b-nvfp4-rtxpro6000-vllm-tp2-mtp3
- model instance id
- nvidia-qwen3-6-35b-a3b-nvfp4--nvfp4
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- No
engine
- graph mode
- unknown
- name
- vllm
- version
- 0.23.0; image build 91df0fad4dc98a67c7659d9dbd915245d5c43d96
serving
- tensor parallel
- 2
Provenance & metadata (3)
facts
- serving.kv cache tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max concurrency · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max context tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- speed sweep ids · provenance · captured at
- 2026-08-27T06:04:13.773Z
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |