Recipe

qwen3-8-27b-ud-q4-k-xl-rx-7900-xtx-24gb-llama-cpp-tp1

qwen3-8-27b-ud-q4-k-xl-rx-7900-xtx-24gb-llama-cpp-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
dflash2-5ecbe1ac1
Accelerators
1
Tensor parallel
1
Context tokens
180,224
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Repository
unsloth/Qwen3.8-27B-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmt1td0t603w5mv01xezz0xz2

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. /home/srid/src/llama.cpp-dflash2/build/bin/llama-server
  2. --model
  3. /home/srid/models/qwen3.8-27b/Qwen3.8-27B-UD-Q4_K_XL.gguf
  4. --mmproj
  5. /home/srid/models/qwen3.8-27b/mmproj-Q8_0-ggmlorg.gguf
  6. --image-min-tokens
  7. 1024
  8. --image-max-tokens
  9. 2048
  10. --host
  11. 127.0.0.1
  12. --port
  13. 5800
  14. --alias
  15. Qwen3.8-27B-Q4
  16. --ctx-size
  17. 180224
  18. --batch-size
  19. 1024
  20. --ubatch-size
  21. 512
  22. --cache-type-k
  23. q8_0
  24. --cache-type-v
  25. q4_0
  26. --flash-attn
  27. on
  28. --n-gpu-layers
  29. 999
  30. --jinja
  31. -md
  32. /home/srid/models/qwen3.8-27b/Qwen3.8-27B-DFlash2-Q4_K_M.gguf
  33. --spec-type
  34. draft-dflash
  35. --spec-draft-n-max
  36. 4
  37. -ngld
  38. 999
  39. --temp
  40. 1.0
  41. --top-p
  42. 0.95
  43. --top-k
  44. 20
  45. --min-p
  46. 0.0
  47. --presence-penalty
  48. 0.0
  49. --repeat-penalty
  50. 1.0
  51. --parallel
  52. 1
  53. --cache-ram
  54. 12288
  55. --cache-idle-slots
  56. --reasoning-budget
  57. 24576
  58. --reasoning-budget-message
  59. Reasoning
  60. budget
  61. exhausted.
  62. Stop
  63. thinking
  64. now
  65. and
  66. immediately
  67. emit
  68. the
  69. tool
  70. call
  71. or
  72. the
  73. final
  74. answer.
  75. --threads-batch
  76. 12
  77. --log-file
  78. /home/srid/.local/share/llm-tuner/llama-server.log
  79. --log-timestamps
FlagValue
--model/home/srid/models/qwen3.8-27b/Qwen3.8-27B-UD-Q4_K_XL.gguf
--mmproj/home/srid/models/qwen3.8-27b/mmproj-Q8_0-ggmlorg.gguf
--image-min-tokens1024
--image-max-tokens2048
--host127.0.0.1
--port5800
--aliasQwen3.8-27B-Q4
--ctx-size180224
--batch-size1024
--ubatch-size512
--cache-type-kq8_0
--cache-type-vq4_0
--flash-attnon
--n-gpu-layers999
-md/home/srid/models/qwen3.8-27b/Qwen3.8-27B-DFlash2-Q4_K_M.gguf
--spec-typedraft-dflash
--spec-draft-n-max4
-ngld999
--temp1.0
--top-p0.95
--top-k20
--min-p0.0
--presence-penalty0.0
--repeat-penalty1.0
--parallel1
--cache-ram12288
--reasoning-budget24576
--reasoning-budget-messageReasoning
--threads-batch12
--log-file/home/srid/.local/share/llm-tuner/llama-server.log

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1180,224191.265.2386.9observedqwen3-8-27b-ud-q4-k-xl-rx-7900-xtx-24gb-llama-cpp-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rx-7900-xtx-24gb
id
qwen3-8-27b-ud-q4-k-xl-rx-7900-xtx-24gb-llama-cpp-tp1
model instance id
unsloth-qwen3-8-27b-gguf--ud-q4-k-xl
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-8-27b-ud-q4-k-xl-rx-7900-xtx-24gb-llama-cpp-tp1-sweep
status
candidate

capabilities

engine

name
llama.cpp
version
dflash2-5ecbe1ac1

serving

max concurrency
1
max context tokens
180,224
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
vulkan
localmaxxing · hardware label
RX 7900 XTX
localmaxxing · notes
llama.cpp fork 'dflash2' on Vulkan (RX 7900 XTX 24GB). Qwen3.8-27B UD-Q4_K_XL + DFlash2 draft model (Q4_K_M), spec-type draft-dflash, draft-n-max 4 (block_size 8 fixed in checkpoint). The GGUF's native MTP head is NOT active (only draft-dflash runs; the alias name 'mtp-n3' is legacy from the old draft-mtp champion). Thinking xhigh enabled (reasoning-budget 24576). KV cache k=q8_0 v=q4_0, flash-attn on, ctx 180224, parallel 1, batch 1024 / ubatch 512, KV RAM 12GB. Server freshly restarted before the run (no prior context). Measured via OpenAI-compatible endpoint through llama-swap proxy: 1 untimed warmup + 3 timed iterations x 256 output tokens; tokSOut = median inter-token streaming rate.
localmaxxing · observed command
/home/srid/src/llama.cpp-dflash2/build/bin/llama-server --model /home/srid/models/qwen3.8-27b/Qwen3.8-27B-UD-Q4_K_XL.gguf --mmproj /home/srid/models/qwen3.8-27b/mmproj-Q8_0-ggmlorg.gguf --image-min-tokens 1024 --image-max-tokens 2048 --host 127.0.0.1 --port 5800 --alias Qwen3.8-27B-Q4 --ctx-size 180224 --batch-size 1024 --ubatch-size 512 --cache-type-k q8_0 --cache-type-v q4_0 --flash-attn on --n-gpu-layers 999 --jinja -md /home/srid/models/qwen3.8-27b/Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -ngld 999 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --parallel 1 --cache-ram 12288 --cache-idle-slots --reasoning-budget 24576 --reasoning-budget-message Reasoning budget exhausted. Stop thinking now and immediately emit the tool call or the final answer. --threads-batch 12 --log-file /home/srid/.local/share/llm-tuner/llama-server.log --log-timestamps
localmaxxing · run id
cmt1td0t603w5mv01xezz0xz2
localmaxxing · tokenized · arguments
/home/srid/src/llama.cpp-dflash2/build/bin/llama-server, --model, /home/srid/models/qwen3.8-27b/Qwen3.8-27B-UD-Q4_K_XL.gguf, --mmproj, /home/srid/models/qwen3.8-27b/mmproj-Q8_0-ggmlorg.gguf, --image-min-tokens, 1024, --image-max-tokens, 2048, --host, 127.0.0.1, --port, 5800, --alias, Qwen3.8-27B-Q4, --ctx-size, 180224, --batch-size, 1024, --ubatch-size, 512, --cache-type-k, q8_0, --cache-type-v, q4_0, --flash-attn, on, --n-gpu-layers, 999, --jinja, -md, /home/srid/models/qwen3.8-27b/Qwen3.8-27B-DFlash2-Q4_K_M.gguf, --spec-type, draft-dflash, --spec-draft-n-max, 4, -ngld, 999, --temp, 1.0, --top-p, 0.95, --top-k, 20, --min-p, 0.0, --presence-penalty, 0.0, --repeat-penalty, 1.0, --parallel, 1, --cache-ram, 12288, --cache-idle-slots, --reasoning-budget, 24576, --reasoning-budget-message, Reasoning, budget, exhausted., Stop, thinking, now, and, immediately, emit, the, tool, call, or, the, final, answer., --threads-batch, 12, --log-file, /home/srid/.local/share/llm-tuner/llama-server.log, --log-timestamps
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmt1td0t603w5mv01xezz0xz2