Recipe

deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4

deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Accelerators
4
Tensor parallel
4
Context tokens
2,048
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Repository
deepseek-ai/DeepSeek-V4-Flash-0731
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmsap6owb00jupm011x6canpd

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

    1. docker
    2. compose
    3. up
    4. -d
    5. #
    6. image
    7. riverhouse/glm52-hybrid:v1.4-cfe43b0
    8. (vLLM
    9. 0.11.2.dev279+eldritch.enlightenment...b12x7bfc945.cu132,
    10. b12x/sparkinfer
    11. 0.30.0)
    1. vllm
    2. serve
    3. /model
    4. --served-model-name
    5. DeepSeek-V4-Flash-0731
    6. --trust-remote-code
    7. --tensor-parallel-size
    8. 4
    9. --kv-cache-dtype
    10. fp8_ds_mla
    11. --moe-backend
    12. b12x
    13. --load-format
    14. safetensors
    15. -cc.pass_config.fuse_allreduce_rms=True
    16. --gpu-memory-utilization
    17. 0.90
    18. --max-model-len
    19. 262144
    20. --max-num-seqs
    21. 8
    22. --max-num-batched-tokens
    23. 8192
    24. --max-cudagraph-capture-size
    25. 48
    26. --async-scheduling
    27. --enable-chunked-prefill
    28. --enable-auto-tool-choice
    29. --tokenizer-mode
    30. deepseek_v4
    31. --tool-call-parser
    32. deepseek_v4
    33. --reasoning-parser
    34. deepseek_v4
    35. --speculative-config
    36. {"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}
FlagValue
--served-model-nameDeepSeek-V4-Flash-0731
--tensor-parallel-size4
--kv-cache-dtypefp8_ds_mla
--moe-backendb12x
--load-formatsafetensors
-cc.pass_config.fuse_allreduce_rmsTrue
--gpu-memory-utilization0.90
--max-model-len262144
--max-num-seqs8
--max-num-batched-tokens8192
--max-cudagraph-capture-size48
--tokenizer-modedeepseek_v4
--tool-call-parserdeepseek_v4
--reasoning-parserdeepseek_v4
--speculative-config{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}

Environment

VariableValue
B12X_DENSE_SPLITK_TURBO1
B12X_MLA_SM120_PREFILL_F16_ROPE1
CUTE_DSL_ARCHsm_120a
NCCL_IB_DISABLE1
NCCL_P2P_DISABLE1
NCCL_P2P_LEVELSYS
NCCL_PROTOLL,LL128,Simple
VLLM_ENABLE_PCIE_ALLREDUCE0
VLLM_USE_B12X_FP8_GEMM0
VLLM_USE_B12X_MOE1
VLLM_USE_B12X_SPARSE_INDEXER1
VLLM_USE_V2_MODEL_RUNNER1

Source notes

Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
12,048497.9231.550.2observeddeepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4
model instance id
deepseek-ai-deepseek-v4-flash-0731--mxfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
--model, deepseek-ai/DeepSeek-V4-Flash-0731, --tensor-parallel-size, 4, --host, 0.0.0.0, --port, 8000, --max-model-len, 2048
container port
8,000
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm

serving

max concurrency
1
max context tokens
2,048
tensor parallel
4
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
RTX PRO 6000 Blackwell
localmaxxing · observed command
docker compose up -d # image riverhouse/glm52-hybrid:v1.4-cfe43b0 (vLLM 0.11.2.dev279+eldritch.enlightenment...b12x7bfc945.cu132, b12x/sparkinfer 0.30.0) vllm serve /model --served-model-name DeepSeek-V4-Flash-0731 --trust-remote-code --tensor-parallel-size 4 --kv-cache-dtype fp8_ds_mla --moe-backend b12x --load-format safetensors -cc.pass_config.fuse_allreduce_rms=True --gpu-memory-utilization 0.90 --max-model-len 262144 --max-num-seqs 8 --max-num-batched-tokens 8192 --max-cudagraph-capture-size 48 --async-scheduling --enable-chunked-prefill --enable-auto-tool-choice --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 --reasoning-parser deepseek_v4 --speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}' # env: CUTE_DSL_ARCH=sm_120a VLLM_USE_V2_MODEL_RUNNER=1 VLLM_USE_B12X_MOE=1 VLLM_USE_B12X_SPARSE_INDEXER=1 VLLM_USE_B12X_FP8_GEMM=0 B12X_DENSE_SPLITK_TURBO=1 B12X_MLA_SM120_PREFILL_F16_ROPE=1 NCCL_P2P_DISABLE=1 NCCL_IB_DISABLE=1 NCCL_P2P_LEVEL=SYS NCCL_PROTO=LL,LL128,Simple VLLM_ENABLE_PCIE_ALLREDUCE=0 # Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.
localmaxxing · run id
cmsap6owb00jupm011x6canpd
localmaxxing · tokenized · arguments
vllm, serve, /model, --served-model-name, DeepSeek-V4-Flash-0731, --trust-remote-code, --tensor-parallel-size, 4, --kv-cache-dtype, fp8_ds_mla, --moe-backend, b12x, --load-format, safetensors, -cc.pass_config.fuse_allreduce_rms=True, --gpu-memory-utilization, 0.90, --max-model-len, 262144, --max-num-seqs, 8, --max-num-batched-tokens, 8192, --max-cudagraph-capture-size, 48, --async-scheduling, --enable-chunked-prefill, --enable-auto-tool-choice, --tokenizer-mode, deepseek_v4, --tool-call-parser, deepseek_v4, --reasoning-parser, deepseek_v4, --speculative-config, {"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}
localmaxxing · tokenized · environment · B12X DENSE SPLITK TURBO
1
localmaxxing · tokenized · environment · B12X MLA SM120 PREFILL F16 ROPE
1
localmaxxing · tokenized · environment · CUTE DSL ARCH
sm_120a
localmaxxing · tokenized · environment · NCCL IB DISABLE
1
localmaxxing · tokenized · environment · NCCL P2P DISABLE
1
localmaxxing · tokenized · environment · NCCL P2P LEVEL
SYS
localmaxxing · tokenized · environment · NCCL PROTO
LL,LL128,Simple
localmaxxing · tokenized · environment · VLLM ENABLE PCIE ALLREDUCE
0
localmaxxing · tokenized · environment · VLLM USE B12X FP8 GEMM
0
localmaxxing · tokenized · environment · VLLM USE B12X MOE
1
localmaxxing · tokenized · environment · VLLM USE B12X SPARSE INDEXER
1
localmaxxing · tokenized · environment · VLLM USE V2 MODEL RUNNER
1
localmaxxing · tokenized · fidelity
faithful
localmaxxing · tokenized · notes
Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.

localmaxxing · tokenized · steps

01234567
dockercomposeup-d#imageriverhouse/glm52-hybrid:v1.4-cfe43b0(vLLM
vllmserve/model--served-model-nameDeepSeek-V4-Flash-0731--trust-remote-code--tensor-parallel-size4

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmsap6owb00jupm011x6canpd