Recipe

deepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1

deepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1

Source-backed single-DGX-Spark K160 REAP profile with a 200K context claim, FP8 KV, MTP2, and full/piecewise CUDA graphs. The archived source's optional eager branch is deliberately absent. CLI launch stays blocked because the validated image is a host-local build rather than a published immutable artifact.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.1.dev17016+g27fd665bd.d20260526
Graph
full-and-piecewise
Accelerators
1
Tensor parallel
1
Context tokens
200,000
Max concurrency
1
KV cache tokens
537,516
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/0xSero/DeepSeek-V4-Flash-180B
Repository
0xSero/DeepSeek-V4-Flash-180B
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0
KV_CACHE_DTYPEfp8
VLLM_ENABLE_DEEPSEEK_V4_SPARSE_MLA_WARMUP0
VLLM_TRITON_MLA_SPARSE1
VLLM_TRITON_MLA_SPARSE_ALLOW_CUDAGRAPH1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1136,534550.133.3248,216.9historicaldeepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1-sweep
1186,390514.124.4362,573.2historicaldeepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1-sweep
1182,112514.718.9353,798.8historicaldeepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
deepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1
model instance id
0xsero-deepseek-v4-flash-180b--nvfp4-mxfp4-experts
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-spark-200k-nvfp4-mxfp4-dgxspark-vllm-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
0.1.dev17016+g27fd665bd.d20260526

serving

kv cache tokens
537,516
max concurrency
1
max context tokens
200,000
tensor parallel
1
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry