Recipe

deepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2

deepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2

Historically accepted DeepSeek-V4-Flash-DSpark NVFP4 across two DGX Spark nodes with TP2, DSpark speculative decoding, NVFP4 DS-MLA KV, and CUDA graphs; awaiting fresh two-node acceptance

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.25.2+anemll-dspark-gx10-0.1.1
Graph
full-and-piecewise
Accelerators
2
Tensor parallel
2
Max concurrency
6
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-DSpark
Repository
deepseek-ai/DeepSeek-V4-Flash-DSpark
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8888

Environment

VariableValue
CUTE_DSL_ARCHsm_121a
FLASHINFER_CUDA_ARCH_LIST12.1a
GPU_MEMORY_UTILIZATION0.80
KV_CACHE_DTYPEnvfp4_ds_mla
MAX_MODEL_LEN1048576
MAX_NUM_BATCHED_TOKENS8192
MAX_NUM_SEQS6
MTP_NUM_TOKENS5
NCCL_IB_DISABLE0
NCCL_NETIB
TORCH_CUDA_ARCH_LIST12.1a
VLLM_USE_B12X_MOE1
VLLM_USE_FLASHINFER_SAMPLER1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
166.6historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
247.2historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
331.9historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
432.8historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
525.7historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
626.8historicaldeepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
dgx-spark-gb10-128gb
id
deepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2
model instance id
deepseek-ai-deepseek-v4-flash-dspark--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-dspark-nvfp4-dgxspark-vllm-tp2-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
0.25.2+anemll-dspark-gx10-0.1.1

serving

max concurrency
6
tensor parallel
2
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry