Recipe

qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1

qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
0.27.2rc1.dev151+g95dc96d1d.d20260827.precompiled
Accelerators
1
Tensor parallel
1
Context tokens
262,144
Max concurrency
16
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4
Repository
RadixArk/Qwen3.8-Flash-Next-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmtbkr6yn002nqq01x9dyaupm

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. export
  2. VLLM_PLE_CPU_OFFLOAD=1
  3. VLLM_PLE_FP8_CHECKPOINT=1
  4. VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200
  5. VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so"
  6. PYTORCH_ALLOC_CONF=expandable_segments:False
  7. MAX_JOBS=15
  8. TORCHINDUCTOR_COMPILE_THREADS=15
  9. OMP_NUM_THREADS=1
  10. LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}";
  11. exec
  12. "$ROOT/.venv-vllm/bin/vllm"
  13. serve
  14. "$ROOT/models/Qwen3.8-Flash-Next-NVFP4"
  15. --served-model-name
  16. Qwen/Qwen3.8-Flash-Next
  17. --host
  18. 0.0.0.0
  19. --port
  20. 8000
  21. --max-model-len
  22. 262144
  23. --gpu-memory-utilization
  24. 0.96
  25. --tensor-parallel-size
  26. 1
  27. --distributed-executor-backend
  28. mp
  29. --max-num-seqs
  30. 16
  31. --max-num-batched-tokens
  32. 8192
  33. --kv-cache-dtype
  34. auto
  35. --enable-prefix-caching
  36. --no-enable-flashinfer-autotune
  37. --moe-backend
  38. humming
  39. --speculative-config
  40. "{\"
  41. method\":\"mtp\",\"num_speculative_tokens\":3}"
  42. --enable-auto-tool-choice
  43. --tool-call-parser
  44. qwen3_coder
  45. --reasoning-parser
  46. qwen3
FlagValue
--served-model-nameQwen/Qwen3.8-Flash-Next
--host0.0.0.0
--port8000
--max-model-len262144
--gpu-memory-utilization0.96
--tensor-parallel-size1
--distributed-executor-backendmp
--max-num-seqs16
--max-num-batched-tokens8192
--kv-cache-dtypeauto
--moe-backendhumming
--speculative-config"{\"
--tool-call-parserqwen3_coder
--reasoning-parserqwen3

Environment

VariableValue
ROOT../qwen38;

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1262,1441,091.3161.268.7observedqwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-pro-6000-blackwell-96gb
id
qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1
model instance id
radixark-qwen3-8-flash-next-nvfp4--nvfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-8-flash-next-nvfp4-rtx-pro-6000-blackwell-96gb-vllm-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
serve, --model, RadixArk/Qwen3.8-Flash-Next-NVFP4, --tensor-parallel-size, 1, --host, 0.0.0.0, --port, 8000, --max-model-len, 262144
container port
8,000
entrypoint
vllm
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm
version
0.27.2rc1.dev151+g95dc96d1d.d20260827.precompiled

serving

max concurrency
16
max context tokens
262,144
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
draft launch.entrypoint · provenance · captured at
2026-09-01T01:41:26Z
draft launch.entrypoint · reason linux-amd64-container-config-entrypoint
draft launch.entrypoint · state known
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-derived-from-source-evidence
serving.max concurrency · state known

metadata

localmaxxing · backend
cuda
localmaxxing · batch size
1
localmaxxing · hardware label
RTX PRO 6000 Blackwell
localmaxxing · notes
Patched vLLM with FP8 PLE checkpoint loading and CPU offload plus a fused sigmoid-GDN kernel, keeping the active weights on the RTX PRO 6000 and the 51B PLE table in system RAM. BF16 KV cache, Humming MoE, fused GDN, and native MTP3. Five 1,024-token single-stream runs achieved median speeds of 161.2 inter-token tok/s and 171.3 client-wall tok/s, with 68.72 ms TTFT. Peak aggregate throughput was 884.6 tok/s across 16 active requests. vLLM commit: 95dc96d1d012a25ff5c3823a1e77197c8dae4654.
localmaxxing · observed command
ROOT=../qwen38; export VLLM_PLE_CPU_OFFLOAD=1 VLLM_PLE_FP8_CHECKPOINT=1 VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200 VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so" PYTORCH_ALLOC_CONF=expandable_segments:False MAX_JOBS=15 TORCHINDUCTOR_COMPILE_THREADS=15 OMP_NUM_THREADS=1 LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}"; exec "$ROOT/.venv-vllm/bin/vllm" serve "$ROOT/models/Qwen3.8-Flash-Next-NVFP4" --served-model-name Qwen/Qwen3.8-Flash-Next --host 0.0.0.0 --port 8000 --max-model-len 262144 --gpu-memory-utilization 0.96 --tensor-parallel-size 1 --distributed-executor-backend mp --max-num-seqs 16 --max-num-batched-tokens 8192 --kv-cache-dtype auto --enable-prefix-caching --no-enable-flashinfer-autotune --moe-backend humming --speculative-config "{\"method\":\"mtp\",\"num_speculative_tokens\":3}" --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
localmaxxing · run id
cmtbkr6yn002nqq01x9dyaupm
localmaxxing · tokenized · arguments
export, VLLM_PLE_CPU_OFFLOAD=1, VLLM_PLE_FP8_CHECKPOINT=1, VLLM_PLE_OFFLOAD_READY_TIMEOUT=1200, VLLM_QWEN38_GDN_PATCH_LIB="$ROOT/kernel_patch/build/qwen38_gdn_sigmoid_patch.so", PYTORCH_ALLOC_CONF=expandable_segments:False, MAX_JOBS=15, TORCHINDUCTOR_COMPILE_THREADS=15, OMP_NUM_THREADS=1, LD_LIBRARY_PATH="$ROOT/.venv-vllm/lib/python3.12/site-packages/nvidia/cu13/lib:${LD_LIBRARY_PATH:-}";, exec, "$ROOT/.venv-vllm/bin/vllm", serve, "$ROOT/models/Qwen3.8-Flash-Next-NVFP4", --served-model-name, Qwen/Qwen3.8-Flash-Next, --host, 0.0.0.0, --port, 8000, --max-model-len, 262144, --gpu-memory-utilization, 0.96, --tensor-parallel-size, 1, --distributed-executor-backend, mp, --max-num-seqs, 16, --max-num-batched-tokens, 8192, --kv-cache-dtype, auto, --enable-prefix-caching, --no-enable-flashinfer-autotune, --moe-backend, humming, --speculative-config, "{\", method\":\"mtp\",\"num_speculative_tokens\":3}", --enable-auto-tool-choice, --tool-call-parser, qwen3_coder, --reasoning-parser, qwen3
localmaxxing · tokenized · environment · ROOT
../qwen38;
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmtbkr6yn002nqq01x9dyaupm