Recipe
deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4
deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4Controller-backed candidate with recovered immutable model, image, vLLM, B12X, and FlashInfer identities. The saved server configuration reports full/piecewise CUDA graphs and 1M context, but no raw client C1/C2/C4 sweep survives, so performance remains unpromoted.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- vLLM e72ad0057b5f382abb40bec1f9d731f6a22600d9; B12X 57422ad9042773165fa0d4e71ae6442fc26b48c9; FlashInfer 25dd814e03791e370f96c3148242f0dc8de504ac
- Graph
- full-and-piecewise
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 1,048,576
- Max concurrency
- 5
- KV cache tokens
- 5,345,692
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731- Repository
- deepseek-ai/DeepSeek-V4-Flash-0731
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Controller configuration
vllm
No container · candidate · controller
Candidate evidence — not a Run contract
Environment
| Variable | Value |
|---|---|
B12X_DENSE_SPLITK_TURBO | 1 |
B12X_MLA_SM120_PREFILL_F16_ROPE | 1 |
CUDA_VISIBLE_DEVICES | 0,1,2,3 |
CUTE_DSL_ARCH | sm_120a |
HF_HUB_OFFLINE | 1 |
NCCL_IB_DISABLE | 1 |
NCCL_P2P_DISABLE | 1 |
NCCL_P2P_LEVEL | SYS |
NCCL_PROTO | LL,LL128,Simple |
NVIDIA_VISIBLE_DEVICES | GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853,GPU-fa982c97-64af-db6a-2ffb-08380e1f9375 |
TRANSFORMERS_OFFLINE | 1 |
VLLM_CACHE_ROOT | /cache |
VLLM_ENABLE_PCIE_ALLREDUCE | 0 |
VLLM_USE_B12X_FP8_GEMM | 0 |
VLLM_USE_B12X_MOE | 1 |
VLLM_USE_B12X_SPARSE_INDEXER | 1 |
VLLM_USE_V2_MODEL_RUNNER | 1 |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- deepseek-v4-flash-0731-fp8-rtxpro6000-vllm-tp4
- model instance id
- deepseek-ai-deepseek-v4-flash-0731--fp8
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- No
engine
- graph mode
- full-and-piecewise
- name
- vllm
- version
- vLLM e72ad0057b5f382abb40bec1f9d731f6a22600d9; B12X 57422ad9042773165fa0d4e71ae6442fc26b48c9; FlashInfer 25dd814e03791e370f96c3148242f0dc8de504ac
serving
- kv cache tokens
- 5,345,692
- max concurrency
- 5
- max context tokens
- 1,048,576
- tensor parallel
- 4
Provenance & metadata (3)
facts
- speed sweep ids · provenance · captured at
- 2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |