Recipe

deepseek-v4-flash-0731-mxfp4-rtx-5090-32gb-llama-cpp-tp1

deepseek-v4-flash-0731-mxfp4-rtx-5090-32gb-llama-cpp-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
4821 (c46ffaa5)
Accelerators
1
Tensor parallel
1
Context tokens
196,608
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/Rabbit-Hole-Ai/DeepSeek-V4-Flash-0731-MXFP4-GGUF
Repository
Rabbit-Hole-Ai/DeepSeek-V4-Flash-0731-MXFP4-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmsy6nels0csbms01n5ob921c

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. llama-server.exe
  2. --host
  3. 127.0.0.1
  4. --port
  5. 8099
  6. -m
  7. D:\\dsv4f-0731-trunk-q8rest-mxfp4moe.gguf
  8. --swa-compress
  9. --ctx-checkpoints
  10. 2
  11. --ctx-checkpoints-tolerance
  12. 5
  13. --ctx-checkpoints-interval
  14. 0
  15. --jinja
  16. --chat-template-file
  17. C:\\s41_template_0731_fixed.jinja
  18. -t
  19. 24
  20. --parallel
  21. 1
  22. --temp
  23. 0.0
  24. --top-k
  25. 1
  26. -c
  27. 196608
  28. -fa
  29. 1
  30. -ctk
  31. f16
  32. -ctv
  33. f16
  34. -ngl
  35. 99
  36. --n-cpu-moe
  37. 39
  38. -b
  39. 7168
  40. -ub
  41. 7168
  42. --rope-scaling
  43. yarn
  44. --rope-scale
  45. 16
  46. --yarn-orig-ctx
  47. 65536
  48. --yarn-beta-fast
  49. 32
  50. --yarn-beta-slow
  51. 1
  52. --spec-type
  53. none
FlagValue
--host127.0.0.1
--port8099
-mD:\\dsv4f-0731-trunk-q8rest-mxfp4moe.gguf
--ctx-checkpoints2
--ctx-checkpoints-tolerance5
--ctx-checkpoints-interval0
--chat-template-fileC:\\s41_template_0731_fixed.jinja
-t24
--parallel1
--temp0.0
--top-k1
-c196608
-fa1
-ctkf16
-ctvf16
-ngl99
--n-cpu-moe39
-b7168
-ub7168
--rope-scalingyarn
--rope-scale16
--yarn-orig-ctx65536
--yarn-beta-fast32
--yarn-beta-slow1
--spec-typenone

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1196,608174.526.43,054.9observeddeepseek-v4-flash-0731-mxfp4-rtx-5090-32gb-llama-cpp-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-5090-32gb
id
deepseek-v4-flash-0731-mxfp4-rtx-5090-32gb-llama-cpp-tp1
model instance id
rabbit-hole-ai-deepseek-v4-flash-0731-mxfp4-gguf--mxfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
deepseek-v4-flash-0731-mxfp4-rtx-5090-32gb-llama-cpp-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
-hf, Rabbit-Hole-Ai/DeepSeek-V4-Flash-0731-MXFP4-GGUF, --n-gpu-layers, 999, --host, 0.0.0.0, --port, 8080, -c, 196608
container port
8,080
environment · LLAMA CACHE
/root/.cache/huggingface
host port
8,080
image
ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107f0f0857e259672d1cb85b71c2
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
gemma-4-12b-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
synthesized · template
llama-cpp-server-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
llama.cpp
version
4821 (c46ffaa5)

serving

max concurrency
1
max context tokens
196,608
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
RTX 5090
localmaxxing · notes
x.com/Rabbit_Hole_Aias-served%20llama-swap%20serving%20flags%20EXCEPT%20sampling%20forced%20greedy%20(temp%200%20top-k%201;%20production%20serves%20temp%201.0);%20median%20of%203%20single-stream%20runs%20on%203%20distinct%20~420-token%20prompts,%20512%20tok%20out,%20each%20run%20first%20sight%20of%20its%20prompt%20(cache_prompt=false,%20no%20prompt%20reuse,%20fresh%20server);%20llama-bench%20baseline%20uses%20ik%20llama-bench%20defaults%20(no%20-c/--swa-compress/--yarn%20support%20in%20that%20binary);%20no-spec%20llama-bench%20baseline:%20pp512=295.853956%20tg128=27.263016;%20context%20allocated%20196608,%20throughput%20measured%20at%20~420-token%20depth;%20peak%20VRAM%20=%20whole-card%20nvidia-smi%20incl%20605%20MiB%20desktop%20baseline;%20raw%20/completion%20(chat%20template%20not%20applied);%20GPU%20power%20=%20max%200.5s%20nvidia-smi%20sample%20during%20generation;%20DeepSeek-V4-Flash-0731%20(in-house%20trunk%20quant:%20https://huggingface.co/Rabbit-Hole-Ai/DeepSeek-V4-Flash-0731-MXFP4-GGUF)
localmaxxing · observed command
llama-server.exe --host 127.0.0.1 --port 8099 -m D:\\dsv4f-0731-trunk-q8rest-mxfp4moe.gguf --swa-compress --ctx-checkpoints 2 --ctx-checkpoints-tolerance 5 --ctx-checkpoints-interval 0 --jinja --chat-template-file C:\\s41_template_0731_fixed.jinja -t 24 --parallel 1 --temp 0.0 --top-k 1 -c 196608 -fa 1 -ctk f16 -ctv f16 -ngl 99 --n-cpu-moe 39 -b 7168 -ub 7168 --rope-scaling yarn --rope-scale 16 --yarn-orig-ctx 65536 --yarn-beta-fast 32 --yarn-beta-slow 1 --spec-type none
localmaxxing · run id
cmsy6nels0csbms01n5ob921c
localmaxxing · tokenized · arguments
llama-server.exe, --host, 127.0.0.1, --port, 8099, -m, D:\\dsv4f-0731-trunk-q8rest-mxfp4moe.gguf, --swa-compress, --ctx-checkpoints, 2, --ctx-checkpoints-tolerance, 5, --ctx-checkpoints-interval, 0, --jinja, --chat-template-file, C:\\s41_template_0731_fixed.jinja, -t, 24, --parallel, 1, --temp, 0.0, --top-k, 1, -c, 196608, -fa, 1, -ctk, f16, -ctv, f16, -ngl, 99, --n-cpu-moe, 39, -b, 7168, -ub, 7168, --rope-scaling, yarn, --rope-scale, 16, --yarn-orig-ctx, 65536, --yarn-beta-fast, 32, --yarn-beta-slow, 1, --spec-type, none
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmsy6nels0csbms01n5ob921c