Recipe
deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4
deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- vllm
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 2,048
- Max concurrency
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731- Repository
- deepseek-ai/DeepSeek-V4-Flash-0731
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
vllm
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
dockercomposeup-d#imageriverhouse/glm52-hybrid:v1.4-cfe43b0(vLLM0.11.2.dev279+eldritch.enlightenment...b12x7bfc945.cu132,b12x/sparkinfer0.30.0)
vllmserve/model--served-model-nameDeepSeek-V4-Flash-0731--trust-remote-code--tensor-parallel-size4--kv-cache-dtypefp8_ds_mla--moe-backendb12x--load-formatsafetensors-cc.pass_config.fuse_allreduce_rms=True--gpu-memory-utilization0.90--max-model-len262144--max-num-seqs8--max-num-batched-tokens8192--max-cudagraph-capture-size48--async-scheduling--enable-chunked-prefill--enable-auto-tool-choice--tokenizer-modedeepseek_v4--tool-call-parserdeepseek_v4--reasoning-parserdeepseek_v4--speculative-config{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}
| Flag | Value |
|---|---|
--served-model-name | DeepSeek-V4-Flash-0731 |
--tensor-parallel-size | 4 |
--kv-cache-dtype | fp8_ds_mla |
--moe-backend | b12x |
--load-format | safetensors |
-cc.pass_config.fuse_allreduce_rms | True |
--gpu-memory-utilization | 0.90 |
--max-model-len | 262144 |
--max-num-seqs | 8 |
--max-num-batched-tokens | 8192 |
--max-cudagraph-capture-size | 48 |
--tokenizer-mode | deepseek_v4 |
--tool-call-parser | deepseek_v4 |
--reasoning-parser | deepseek_v4 |
--speculative-config | {"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"} |
Environment
| Variable | Value |
|---|---|
B12X_DENSE_SPLITK_TURBO | 1 |
B12X_MLA_SM120_PREFILL_F16_ROPE | 1 |
CUTE_DSL_ARCH | sm_120a |
NCCL_IB_DISABLE | 1 |
NCCL_P2P_DISABLE | 1 |
NCCL_P2P_LEVEL | SYS |
NCCL_PROTO | LL,LL128,Simple |
VLLM_ENABLE_PCIE_ALLREDUCE | 0 |
VLLM_USE_B12X_FP8_GEMM | 0 |
VLLM_USE_B12X_MOE | 1 |
VLLM_USE_B12X_SPARSE_INDEXER | 1 |
VLLM_USE_V2_MODEL_RUNNER | 1 |
Source notes
Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 2,048 | 497.9 | 231.5 | 50.2 | observed | deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4
- model instance id
- deepseek-ai-deepseek-v4-flash-0731--mxfp4
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- deepseek-v4-flash-0731-mxfp4-rtx-pro-6000-blackwell-96gb-vllm-tp4-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- --model, deepseek-ai/DeepSeek-V4-Flash-0731, --tensor-parallel-size, 4, --host, 0.0.0.0, --port, 8000, --max-model-len, 2048
- container port
- 8,000
- host port
- 8,000
- image
- vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
- synthesized · template
- vllm-openai-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- vllm
serving
- max concurrency
- 1
- max context tokens
- 2,048
- tensor parallel
- 4
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
metadata
- localmaxxing · backend
- cuda
- localmaxxing · hardware label
- RTX PRO 6000 Blackwell
- localmaxxing · observed command
- docker compose up -d # image riverhouse/glm52-hybrid:v1.4-cfe43b0 (vLLM 0.11.2.dev279+eldritch.enlightenment...b12x7bfc945.cu132, b12x/sparkinfer 0.30.0) vllm serve /model --served-model-name DeepSeek-V4-Flash-0731 --trust-remote-code --tensor-parallel-size 4 --kv-cache-dtype fp8_ds_mla --moe-backend b12x --load-format safetensors -cc.pass_config.fuse_allreduce_rms=True --gpu-memory-utilization 0.90 --max-model-len 262144 --max-num-seqs 8 --max-num-batched-tokens 8192 --max-cudagraph-capture-size 48 --async-scheduling --enable-chunked-prefill --enable-auto-tool-choice --tokenizer-mode deepseek_v4 --tool-call-parser deepseek_v4 --reasoning-parser deepseek_v4 --speculative-config '{"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}' # env: CUTE_DSL_ARCH=sm_120a VLLM_USE_V2_MODEL_RUNNER=1 VLLM_USE_B12X_MOE=1 VLLM_USE_B12X_SPARSE_INDEXER=1 VLLM_USE_B12X_FP8_GEMM=0 B12X_DENSE_SPLITK_TURBO=1 B12X_MLA_SM120_PREFILL_F16_ROPE=1 NCCL_P2P_DISABLE=1 NCCL_IB_DISABLE=1 NCCL_P2P_LEVEL=SYS NCCL_PROTO=LL,LL128,Simple VLLM_ENABLE_PCIE_ALLREDUCE=0 # Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.
- localmaxxing · run id
- cmsap6owb00jupm011x6canpd
- localmaxxing · tokenized · arguments
- vllm, serve, /model, --served-model-name, DeepSeek-V4-Flash-0731, --trust-remote-code, --tensor-parallel-size, 4, --kv-cache-dtype, fp8_ds_mla, --moe-backend, b12x, --load-format, safetensors, -cc.pass_config.fuse_allreduce_rms=True, --gpu-memory-utilization, 0.90, --max-model-len, 262144, --max-num-seqs, 8, --max-num-batched-tokens, 8192, --max-cudagraph-capture-size, 48, --async-scheduling, --enable-chunked-prefill, --enable-auto-tool-choice, --tokenizer-mode, deepseek_v4, --tool-call-parser, deepseek_v4, --reasoning-parser, deepseek_v4, --speculative-config, {"method":"dspark","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}
- localmaxxing · tokenized · environment · B12X DENSE SPLITK TURBO
- 1
- localmaxxing · tokenized · environment · B12X MLA SM120 PREFILL F16 ROPE
- 1
- localmaxxing · tokenized · environment · CUTE DSL ARCH
- sm_120a
- localmaxxing · tokenized · environment · NCCL IB DISABLE
- 1
- localmaxxing · tokenized · environment · NCCL P2P DISABLE
- 1
- localmaxxing · tokenized · environment · NCCL P2P LEVEL
- SYS
- localmaxxing · tokenized · environment · NCCL PROTO
- LL,LL128,Simple
- localmaxxing · tokenized · environment · VLLM ENABLE PCIE ALLREDUCE
- 0
- localmaxxing · tokenized · environment · VLLM USE B12X FP8 GEMM
- 0
- localmaxxing · tokenized · environment · VLLM USE B12X MOE
- 1
- localmaxxing · tokenized · environment · VLLM USE B12X SPARSE INDEXER
- 1
- localmaxxing · tokenized · environment · VLLM USE V2 MODEL RUNNER
- 1
- localmaxxing · tokenized · fidelity
- faithful
- localmaxxing · tokenized · notes
- Notes: Threadripper PRO topology deadlocks on PCIe P2P, so P2P is disabled. b12x MoE is used; b12x MLA (B12X_MLA_SPARSE) is NOT usable for DeepSeek-V4 in b12x 0.30.0 (DSV4 topk-128 MG prefill dispatcher passes scale_format= to io_issue_gather_dsv4_nope(), which rejects it). VLLM_USE_B12X_FP8_GEMM must be 0 with DSpark or deep_gemm fp8_einsum hits a layout assertion during drafter warmup. DSpark requires decode-context-parallel-size 1. DeepseekV4 ds_mla accepts fp8_ds_mla only, not nvfp4_ds_mla.
localmaxxing · tokenized · steps
| 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|
| docker | compose | up | -d | # | image | riverhouse/glm52-hybrid:v1.4-cfe43b0 | (vLLM |
| vllm | serve | /model | --served-model-name | DeepSeek-V4-Flash-0731 | --trust-remote-code | --tensor-parallel-size | 4 |
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmsap6owb00jupm011x6canpd ↗ |