Recipe

deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4

deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4

Controller-backed candidate with recovered immutable model, image, vLLM, B12X, and FlashInfer identities. The saved server configuration reports full/piecewise CUDA graphs and 1M context, but no raw client C1/C2/C4 sweep survives, so performance remains unpromoted.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
vLLM e72ad0057b5f382abb40bec1f9d731f6a22600d9; B12X 57422ad9042773165fa0d4e71ae6442fc26b48c9; FlashInfer 25dd814e03791e370f96c3148242f0dc8de504ac
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
1,048,576
Max concurrency
5
KV cache tokens
5,345,692
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Repository
deepseek-ai/DeepSeek-V4-Flash-0731
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Controller configuration

vllm

No container · candidate · controller

Candidate evidence — not a Run contract

Environment

VariableValue
B12X_DENSE_SPLITK_TURBO1
B12X_MLA_SM120_PREFILL_F16_ROPE1
CUDA_VISIBLE_DEVICES0,1,2,3
CUTE_DSL_ARCHsm_120a
HF_HUB_OFFLINE1
NCCL_IB_DISABLE1
NCCL_P2P_DISABLE1
NCCL_P2P_LEVELSYS
NCCL_PROTOLL,LL128,Simple
NVIDIA_VISIBLE_DEVICESGPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853,GPU-fa982c97-64af-db6a-2ffb-08380e1f9375
TRANSFORMERS_OFFLINE1
VLLM_CACHE_ROOT/cache
VLLM_ENABLE_PCIE_ALLREDUCE0
VLLM_USE_B12X_FP8_GEMM0
VLLM_USE_B12X_MOE1
VLLM_USE_B12X_SPARSE_INDEXER1
VLLM_USE_V2_MODEL_RUNNER1

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4
model instance id
deepseek-ai-deepseek-v4-flash-0731--fp8
recipe source
0xsero
schema version
local-ai-registry/v1
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
vLLM e72ad0057b5f382abb40bec1f9d731f6a22600d9; B12X 57422ad9042773165fa0d4e71ae6442fc26b48c9; FlashInfer 25dd814e03791e370f96c3148242f0dc8de504ac

serving

kv cache tokens
5,345,692
max concurrency
5
max context tokens
1,048,576
tensor parallel
4
Provenance & metadata (3)

facts

speed sweep ids · provenance · captured at
2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry