Recipe

deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4

deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4

Pinned-source DeepSeek V4 Flash FP8 SGLang TP4 candidate using the SM120 sparse-decode extension, built-in DeepSeek V4 encoding, EAGLE2, and CUDA graphs. It is intentionally indexed under SGLang. The tag-only image, unpinned model revision, and host-built extension keep CLI launch blocked.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
deepseek-v4-blackwell image; exact SGLang commit unreported
Graph
full
Accelerators
4
Tensor parallel
4
Context tokens
300,000
Max concurrency
1
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/sgl-project/DeepSeek-V4-Flash-FP8
Repository
sgl-project/DeepSeek-V4-Flash-FP8
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

sglang

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0,1,2,3
PYTHONPATH/dsv4
SGLANG_ENABLE_SPEC_V2True
SGLANG_ENABLE_THINKING1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
18,192701.637.611,337historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
116,3841,2503812,770historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
132,7681,273.135.525,101historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
165,5361,145.130.355,850historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
1131,07294624.1135,271historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
1196,000793.819.3246,863historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
1300,000646.415.1464,072historicaldeepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4
model instance id
sgl-project-deepseek-v4-flash-fp8--fp8
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-fp8-rtxpro6000-sglang-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full
name
sglang
version
deepseek-v4-blackwell image; exact SGLang commit unreported

serving

max concurrency
1
max context tokens
300,000
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry