Recipe

deepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1

deepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1

MiaAI-Lab single-Spark DeepSeek V4 Flash 0731 EXL3 3.0bpw profile with SparkInfer, DSpark K5, native NVFP4 DS-MLA KV, and graph captures

Record

Status
candidate
Source
mialabs
Engine
sparkinfer
Engine version
NVIDIA vLLM 26.02 base
Graph
full-and-piecewise
Accelerators
1
Tensor parallel
1
Context tokens
370,104
Max concurrency
1
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/0xSero/deepseek-v4-flash-0731-spark
Repository
0xSero/deepseek-v4-flash-0731-spark
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Script configuration

sparkinfer

No container · candidate · script

Candidate evidence — not a Run contract

Environment

VariableValue
CUDAGRAPH_CAPTURE_SIZES6,12,24
GPU_MEMORY_UTILIZATION0.94
KV_RECORDstock432
MAX_CUDAGRAPH_CAPTURE_SIZE24
MAX_MODEL_LEN384000
MAX_NUM_BATCHED_TOKENS8224
MAX_NUM_SEQS1
MODEL_REVISION22f28d32b9b29b4352eaa380ff8c2c170b2847ab

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1320,00063047historicaldeepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1-sweep
1370,10462544historicaldeepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
deepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1
model instance id
0xsero-deepseek-v4-flash-0731-spark--3-0-bpw-reap-k216
recipe source
mialabs
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-0731-exl3-3bpw-dgxspark-sparkinfer-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
sparkinfer
version
NVIDIA vLLM 26.02 base

serving

max concurrency
1
max context tokens
370,104
tensor parallel
1
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry