Recipe
deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4
deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- hipfire
- Engine version
- 0.3.0+5ccef2df2719
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 2,052
- Max concurrency
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash- Repository
- deepseek-ai/DeepSeek-V4-Flash
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
hipfire
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
python3scripts/serve_harness.py--modeldeepseek-v4-flash-0731.mq2r--kvf32--kv-backendcontiguous--speculationoff--mtpoff--dflashoff--thinkingoff--samplinggreedy--max-tokens512--modebattery--prompts-filebenchmarks/prompts/ds4-gfx942-ar-2048.txt--devices0,1,2,3--tp4
| Flag | Value |
|---|---|
--model | deepseek-v4-flash-0731.mq2r |
--kv | f32 |
--kv-backend | contiguous |
--speculation | off |
--mtp | off |
--dflash | off |
--thinking | off |
--sampling | greedy |
--max-tokens | 512 |
--mode | battery |
--prompts-file | benchmarks/prompts/ds4-gfx942-ar-2048.txt |
--devices | 0,1,2,3 |
--tp | 4 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 2,052 | 389.2 | 54.3 | 6,095 | observed | deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- radeon-ai-pro-r9700-32gb
- id
- deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4
- model instance id
- deepseek-ai-deepseek-v4-flash--mq2r
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4-sweep
- status
- candidate
capabilities
engine
- name
- hipfire
- version
- 0.3.0+5ccef2df2719
serving
- max concurrency
- 1
- max context tokens
- 2,052
- tensor parallel
- 4
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
metadata
- localmaxxing · backend
- rocm
- localmaxxing · hardware label
- Radeon AI Pro R9700
- localmaxxing · notes
- SINGLE-STREAM. Autoregressive decode, batch 1, greedy, speculation off, tensor-parallel 4 across 4 x Radeon AI PRO R9700 (devices 0,1,2,3). hipfire 0.3.0+5ccef2df2719, backend rocm, gfx1201. quant: MQ2R = MagnumQuant 2-bit Redline (MQ3-Lloyd routed tier); model deepseek-v4-flash-0731.mq2r sha256 cbf2bbcf... KV: explicit f32 compressor cache, verified per run by the loader line 'compressor_cache=f32' with zero F16-fallback lines. DeepSeek-V4 implements only F32 or F16 compressor storage and silently widens sub-F16 selectors up to F16, so a run that asks for q8 actually executes F16; the engine advises 'Pass --kv f32 for the golden configuration'. This row is the golden f32 config. method: scripts/serve_harness.py battery, 3 fresh processes, GPUs verified idle before each launch, median reported. prompt ds4-gfx942-ar-2048.txt md5 25e22fae..., 2052 prompt tokens, 512 generated, finish=length, output coherent. samples - decode 54.370 / 54.266 / 54.264 tok/s; prefill 390.97 / 389.19 / 389.04 tok/s; TTFT 6095 / 6120 / 6086 ms. An identical n=3 battery on the F16 path measures within run-to-run noise of these figures on both TP3 and TP4; f32 is reported because it is the correct configuration, not because it is faster. evidence: 20260810T194943Z-ds4-f32-tp-triplicate not measured: peak VRAM (not instrumented this campaign)
- localmaxxing · observed command
- python3 scripts/serve_harness.py --model deepseek-v4-flash-0731.mq2r --kv f32 --kv-backend contiguous --speculation off --mtp off --dflash off --thinking off --sampling greedy --max-tokens 512 --mode battery --prompts-file benchmarks/prompts/ds4-gfx942-ar-2048.txt --devices 0,1,2,3 --tp 4
- localmaxxing · run id
- cmsnp21x400k5o001hc7cubcl
- localmaxxing · tokenized · arguments
- python3, scripts/serve_harness.py, --model, deepseek-v4-flash-0731.mq2r, --kv, f32, --kv-backend, contiguous, --speculation, off, --mtp, off, --dflash, off, --thinking, off, --sampling, greedy, --max-tokens, 512, --mode, battery, --prompts-file, benchmarks/prompts/ds4-gfx942-ar-2048.txt, --devices, 0,1,2,3, --tp, 4
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmsnp21x400k5o001hc7cubcl ↗ |