Recipe

qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1

qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
hipfire
Engine version
00374bb7dfab0235a25f6ba61cf12c0d711f3703
Accelerators
1
Tensor parallel
1
Context tokens
32,768
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/Qwen/Qwen3.8-27B
Repository
Qwen/Qwen3.8-27B
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

hipfire

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmt62jgz20032lk01uo7i0p1d

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. dflash_spec_demo
  2. --target
  3. q38-fixed.mq4v2.mq4
  4. --draft
  5. qwen38-27b-dflash2-mq4v2.hfq
  6. --prompt-file
  7. merge_sort_thinking_off.txt
  8. --max
  9. 256
  10. --temp
  11. 0
  12. --no-chatml
  13. --kv-mode
  14. q8
  15. --ctx
  16. 32768
  17. --no-adaptive-b
FlagValue
--targetq38-fixed.mq4v2.mq4
--draftqwen38-27b-dflash2-mq4v2.hfq
--prompt-filemerge_sort_thinking_off.txt
--max256
--temp0
--kv-modeq8
--ctx32768

Environment

VariableValue
HIPFIRE_HFQ4G256_LDSSTAGE1
HIPFIRE_VERIFY_GRAPH0
HIP_VISIBLE_DEVICES0

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
32,768435.2286.3115.1observedqwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
radeon-ai-pro-r9700-32gb
id
qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1
model instance id
qwen-qwen3-8-27b--int4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1-sweep
status
candidate

capabilities

engine

name
hipfire
version
00374bb7dfab0235a25f6ba61cf12c0d711f3703

serving

max context tokens
32,768
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown

metadata

localmaxxing · backend
rocm
localmaxxing · batch size
1
localmaxxing · hardware label
Radeon AI Pro R9700
localmaxxing · notes
32K context point in a controlled Qwen3.8-27B R9700 context sweep. R9700 gfx1201 only; Strix Halo/Radeon 8060S iGPU unused. One warmup plus three fresh measured processes with clean driver-baseline waits; median decode/prefill/TTFT/total and maximum measured workload VRAM. tokSTotal = (27 prompt + 157 output tokens) / (prefill_secs + decode_secs), calculated per run before the median. Peak workload VRAM = max(/sys/class/drm/card1/device/mem_info_vram_used) minus a clean 59,936,768-byte baseline, sampled every 50 ms. Exact target is independently quantized MQ4V2 from Qwen/Qwen3.8-27B revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; INT4 is the nearest structured label because MQ4V2 has no dedicated enum. Hipfire DFlash2, Q8 KV, runtime block 16, temp 0. All outputs byte-identical and correct merge-sort code; 157 tokens, tau 13.1818, acceptance 0.8788. The 128K Q8-KV point could not load: hipMalloc OOM after target+draft reached 21.32 GB.
localmaxxing · observed command
HIP_VISIBLE_DEVICES=0 HIPFIRE_HFQ4G256_LDSSTAGE=1 HIPFIRE_VERIFY_GRAPH=0 dflash_spec_demo --target q38-fixed.mq4v2.mq4 --draft qwen38-27b-dflash2-mq4v2.hfq --prompt-file merge_sort_thinking_off.txt --max 256 --temp 0 --no-chatml --kv-mode q8 --ctx 32768 --no-adaptive-b
localmaxxing · run id
cmt62jgz20032lk01uo7i0p1d
localmaxxing · tokenized · arguments
dflash_spec_demo, --target, q38-fixed.mq4v2.mq4, --draft, qwen38-27b-dflash2-mq4v2.hfq, --prompt-file, merge_sort_thinking_off.txt, --max, 256, --temp, 0, --no-chatml, --kv-mode, q8, --ctx, 32768, --no-adaptive-b
localmaxxing · tokenized · environment · HIPFIRE HFQ4G256 LDSSTAGE
1
localmaxxing · tokenized · environment · HIPFIRE VERIFY GRAPH
0
localmaxxing · tokenized · environment · HIP VISIBLE DEVICES
0
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmt62jgz20032lk01uo7i0p1d