Recipe

deepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2

deepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2

MiaAI-Lab official DeepSeek V4 Flash 0731 NVFP4 across two DGX Sparks with TP2, 1M configured context, MTP-5, NVFP4 DS-MLA KV, and CUDA graphs

Record

Status
candidate
Source
mialabs
Engine
vllm
Engine version
0.25.2+anemll-dspark-gx10-0.1.1
Graph
full-and-piecewise
Accelerators
2
Tensor parallel
2
Context tokens
899,994
Max concurrency
6
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Repository
deepseek-ai/DeepSeek-V4-Flash-0731
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Script configuration

vllm

No container · candidate · script

Candidate evidence — not a Run contract

Environment

VariableValue
DSPARK_REVISION9e165c30e2704aec5d9d593cce3eebd58bbef1cb
GPU_MEMORY_UTILIZATION_TEXT0.835
KV_CACHE_DTYPEnvfp4_ds_mla
MAX_MODEL_LEN1048576
MAX_NUM_BATCHED_TOKENS8192
MAX_NUM_SEQS6
MTP_NUM_TOKENS5
VLLM_USE_BREAKABLE_CUDAGRAPH0

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
125644775.4historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
225635758.3historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
425622246.8historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
625619736.9historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
12,0482,56368.8historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
22,0481,91157historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
42,0481,50544historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
62,04834234.7historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
18,1921,71373.9historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
28,1921,17649.8historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
48,19257837.4historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
68,19245423.6historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
132,7681,42864historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
232,7681,28741.5historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
432,76875617.4historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
632,76855010.8historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
1131,0721,66565.2historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
2131,0721,30630.9historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
1899,994874.8historicaldeepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
dgx-spark-gb10-128gb
id
deepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2
model instance id
deepseek-ai-deepseek-v4-flash-0731--nvfp4
recipe source
mialabs
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-0731-nvfp4-dgxspark-vllm-tp2-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
0.25.2+anemll-dspark-gx10-0.1.1

serving

max concurrency
6
max context tokens
899,994
tensor parallel
2
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry