Recipe

deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4

deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
hipfire
Engine version
0.3.0+5ccef2df2719
Accelerators
4
Tensor parallel
4
Context tokens
2,052
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash
Repository
deepseek-ai/DeepSeek-V4-Flash
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

hipfire

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmsnp21x400k5o001hc7cubcl

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. python3
  2. scripts/serve_harness.py
  3. --model
  4. deepseek-v4-flash-0731.mq2r
  5. --kv
  6. f32
  7. --kv-backend
  8. contiguous
  9. --speculation
  10. off
  11. --mtp
  12. off
  13. --dflash
  14. off
  15. --thinking
  16. off
  17. --sampling
  18. greedy
  19. --max-tokens
  20. 512
  21. --mode
  22. battery
  23. --prompts-file
  24. benchmarks/prompts/ds4-gfx942-ar-2048.txt
  25. --devices
  26. 0,1,2,3
  27. --tp
  28. 4
FlagValue
--modeldeepseek-v4-flash-0731.mq2r
--kvf32
--kv-backendcontiguous
--speculationoff
--mtpoff
--dflashoff
--thinkingoff
--samplinggreedy
--max-tokens512
--modebattery
--prompts-filebenchmarks/prompts/ds4-gfx942-ar-2048.txt
--devices0,1,2,3
--tp4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
12,052389.254.36,095observeddeepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
radeon-ai-pro-r9700-32gb
id
deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4
model instance id
deepseek-ai-deepseek-v4-flash--mq2r
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-mq2r-radeon-ai-pro-r9700-32gb-hipfire-tp4-sweep
status
candidate

capabilities

engine

name
hipfire
version
0.3.0+5ccef2df2719

serving

max concurrency
1
max context tokens
2,052
tensor parallel
4
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
rocm
localmaxxing · hardware label
Radeon AI Pro R9700
localmaxxing · notes
SINGLE-STREAM. Autoregressive decode, batch 1, greedy, speculation off, tensor-parallel 4 across 4 x Radeon AI PRO R9700 (devices 0,1,2,3). hipfire 0.3.0+5ccef2df2719, backend rocm, gfx1201. quant: MQ2R = MagnumQuant 2-bit Redline (MQ3-Lloyd routed tier); model deepseek-v4-flash-0731.mq2r sha256 cbf2bbcf... KV: explicit f32 compressor cache, verified per run by the loader line 'compressor_cache=f32' with zero F16-fallback lines. DeepSeek-V4 implements only F32 or F16 compressor storage and silently widens sub-F16 selectors up to F16, so a run that asks for q8 actually executes F16; the engine advises 'Pass --kv f32 for the golden configuration'. This row is the golden f32 config. method: scripts/serve_harness.py battery, 3 fresh processes, GPUs verified idle before each launch, median reported. prompt ds4-gfx942-ar-2048.txt md5 25e22fae..., 2052 prompt tokens, 512 generated, finish=length, output coherent. samples - decode 54.370 / 54.266 / 54.264 tok/s; prefill 390.97 / 389.19 / 389.04 tok/s; TTFT 6095 / 6120 / 6086 ms. An identical n=3 battery on the F16 path measures within run-to-run noise of these figures on both TP3 and TP4; f32 is reported because it is the correct configuration, not because it is faster. evidence: 20260810T194943Z-ds4-f32-tp-triplicate not measured: peak VRAM (not instrumented this campaign)
localmaxxing · observed command
python3 scripts/serve_harness.py --model deepseek-v4-flash-0731.mq2r --kv f32 --kv-backend contiguous --speculation off --mtp off --dflash off --thinking off --sampling greedy --max-tokens 512 --mode battery --prompts-file benchmarks/prompts/ds4-gfx942-ar-2048.txt --devices 0,1,2,3 --tp 4
localmaxxing · run id
cmsnp21x400k5o001hc7cubcl
localmaxxing · tokenized · arguments
python3, scripts/serve_harness.py, --model, deepseek-v4-flash-0731.mq2r, --kv, f32, --kv-backend, contiguous, --speculation, off, --mtp, off, --dflash, off, --thinking, off, --sampling, greedy, --max-tokens, 512, --mode, battery, --prompts-file, benchmarks/prompts/ds4-gfx942-ar-2048.txt, --devices, 0,1,2,3, --tp, 4
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmsnp21x400k5o001hc7cubcl