Recipe
qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1
qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- vllm
- Engine version
- 0.27.2rc1.dev151+g95dc96d1d.d20260827.precompiled
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 262,144
- Max concurrency
- 16
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4- Repository
- RadixArk/Qwen3.8-Flash-Next-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
vllm
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
exportVLLM_PLE_CPU_OFFLOAD=1VLLM_PLE_FP8_CHECKPOINT=1VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so"PYTORCH_ALLOC_CONF=expandable_segments:FalseMAX_JOBS=15TORCHINDUCTOR_COMPILE_THREADS=15OMP_NUM_THREADS=1LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}";exec"$ROOT/.venv-vllm/bin/vllm"serve"$ROOT/models/Qwen3.8-Flash-Next-NVFP4"--served-model-nameQwen/Qwen3.8-Flash-Next--host0.0.0.0--port8000--max-model-len262144--gpu-memory-utilization0.96--tensor-parallel-size1--distributed-executor-backendmp--max-num-seqs16--max-num-batched-tokens8192--kv-cache-dtypeauto--enable-prefix-caching--no-enable-flashinfer-autotune--moe-backendhumming--speculative-config"{\"method\":\"mtp\",\"num_speculative_tokens\":3}"--enable-auto-tool-choice--tool-call-parserqwen3_coder--reasoning-parserqwen3
| Flag | Value |
|---|---|
--served-model-name | Qwen/Qwen3.8-Flash-Next |
--host | 0.0.0.0 |
--port | 8000 |
--max-model-len | 262144 |
--gpu-memory-utilization | 0.96 |
--tensor-parallel-size | 1 |
--distributed-executor-backend | mp |
--max-num-seqs | 16 |
--max-num-batched-tokens | 8192 |
--kv-cache-dtype | auto |
--moe-backend | humming |
--speculative-config | "{\" |
--tool-call-parser | qwen3_coder |
--reasoning-parser | qwen3 |
Environment
| Variable | Value |
|---|---|
ROOT | ../qwen38; |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 262,144 | 1,091.3 | 161.2 | 68.7 | observed | qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1
- model instance id
- radixark-qwen3-8-flash-next-nvfp4--nvfp4
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- serve, --model, RadixArk/Qwen3.8-Flash-Next-NVFP4, --tensor-parallel-size, 1, --host, 0.0.0.0, --port, 8000, --max-model-len, 262144
- container port
- 8,000
- entrypoint
- vllm
- host port
- 8,000
- image
- vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
- synthesized · template
- vllm-openai-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- vllm
- version
- 0.27.2rc1.dev151+g95dc96d1d.d20260827.precompiled
serving
- max concurrency
- 16
- max context tokens
- 262,144
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- draft launch.entrypoint · provenance · captured at
- 2026-09-01T01:41:26Z
draft launch.entrypoint · reason linux-amd64-container-config-entrypoint
draft launch.entrypoint · state known
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max concurrency · provenance · captured at
- 2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-derived-from-source-evidence
serving.max concurrency · state known
metadata
- localmaxxing · backend
- cuda
- localmaxxing · batch size
- 1
- localmaxxing · hardware label
- RTX PRO 6000 Blackwell
- localmaxxing · notes
- Patched vLLM with FP8 PLE checkpoint loading and CPU offload plus a fused sigmoid-GDN kernel, keeping the active weights on the RTX PRO 6000 and the 51B PLE table in system RAM. BF16 KV cache, Humming MoE, fused GDN, and native MTP3. Five 1,024-token single-stream runs achieved median speeds of 161.2 inter-token tok/s and 171.3 client-wall tok/s, with 68.72 ms TTFT. Peak aggregate throughput was 884.6 tok/s across 16 active requests. vLLM commit: 95dc96d1d012a25ff5c3823a1e77197c8dae4654.
- localmaxxing · observed command
- ROOT=../qwen38; export VLLM_PLE_CPU_OFFLOAD=1 VLLM_PLE_FP8_CHECKPOINT=1 VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200 VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so" PYTORCH_ALLOC_CONF=expandable_segments:False MAX_JOBS=15 TORCHINDUCTOR_COMPILE_THREADS=15 OMP_NUM_THREADS=1 LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}"; exec "$ROOT/.venv-vllm/bin/vllm" serve "$ROOT/models/Qwen3.8-Flash-Next-NVFP4" --served-model-name Qwen/Qwen3.8-Flash-Next --host 0.0.0.0 --port 8000 --max-model-len 262144 --gpu-memory-utilization 0.96 --tensor-parallel-size 1 --distributed-executor-backend mp --max-num-seqs 16 --max-num-batched-tokens 8192 --kv-cache-dtype auto --enable-prefix-caching --no-enable-flashinfer-autotune --moe-backend humming --speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":3}" --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
- localmaxxing · run id
- cmtbkr6yn002nqq01x9dyaupm
- localmaxxing · tokenized · arguments
- export, VLLM_PLE_CPU_OFFLOAD=1, VLLM_PLE_FP8_CHECKPOINT=1, VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200, VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so", PYTORCH_ALLOC_CONF=expandable_segments:False, MAX_JOBS=15, TORCHINDUCTOR_COMPILE_THREADS=15, OMP_NUM_THREADS=1, LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}";, exec, "$ROOT/.venv-vllm/bin/vllm", serve, "$ROOT/models/Qwen3.8-Flash-Next-NVFP4", --served-model-name, Qwen/Qwen3.8-Flash-Next, --host, 0.0.0.0, --port, 8000, --max-model-len, 262144, --gpu-memory-utilization, 0.96, --tensor-parallel-size, 1, --distributed-executor-backend, mp, --max-num-seqs, 16, --max-num-batched-tokens, 8192, --kv-cache-dtype, auto, --enable-prefix-caching, --no-enable-flashinfer-autotune, --moe-backend, humming, --speculative-config, "{\", method\":\"mtp\",\"num_speculative_tokens\":3}", --enable-auto-tool-choice, --tool-call-parser, qwen3_coder, --reasoning-parser, qwen3
- localmaxxing · tokenized · environment · ROOT
- ../qwen38;
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmtbkr6yn002nqq01x9dyaupm ↗ |