Recipe
qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1
qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- hipfire
- Engine version
- 00374bb7dfab0235a25f6ba61cf12c0d711f3703
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 32,768
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/Qwen/Qwen3.8-27B- Repository
- Qwen/Qwen3.8-27B
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
hipfire
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
dflash_spec_demo--targetq38-fixed.mq4v2.mq4--draftqwen38-27b-dflash2-mq4v2.hfq--prompt-filemerge_sort_thinking_off.txt--max256--temp0--no-chatml--kv-modeq8--ctx32768--no-adaptive-b
| Flag | Value |
|---|---|
--target | q38-fixed.mq4v2.mq4 |
--draft | qwen38-27b-dflash2-mq4v2.hfq |
--prompt-file | merge_sort_thinking_off.txt |
--max | 256 |
--temp | 0 |
--kv-mode | q8 |
--ctx | 32768 |
Environment
| Variable | Value |
|---|---|
HIPFIRE_HFQ4G256_LDSSTAGE | 1 |
HIPFIRE_VERIFY_GRAPH | 0 |
HIP_VISIBLE_DEVICES | 0 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| — | 32,768 | 435.2 | 286.3 | 115.1 | observed | qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- radeon-ai-pro-r9700-32gb
- id
- qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1
- model instance id
- qwen-qwen3-8-27b--int4
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen3-8-27b-int4-radeon-ai-pro-r9700-32gb-hipfire-tp1-sweep
- status
- candidate
capabilities
engine
- name
- hipfire
- version
- 00374bb7dfab0235a25f6ba61cf12c0d711f3703
serving
- max context tokens
- 32,768
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max concurrency · provenance · captured at
- 2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown
metadata
- localmaxxing · backend
- rocm
- localmaxxing · batch size
- 1
- localmaxxing · hardware label
- Radeon AI Pro R9700
- localmaxxing · notes
- 32K context point in a controlled Qwen3.8-27B R9700 context sweep. R9700 gfx1201 only; Strix Halo/Radeon 8060S iGPU unused. One warmup plus three fresh measured processes with clean driver-baseline waits; median decode/prefill/TTFT/total and maximum measured workload VRAM. tokSTotal = (27 prompt + 157 output tokens) / (prefill_secs + decode_secs), calculated per run before the median. Peak workload VRAM = max(/sys/class/drm/card1/device/mem_info_vram_used) minus a clean 59,936,768-byte baseline, sampled every 50 ms. Exact target is independently quantized MQ4V2 from Qwen/Qwen3.8-27B revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0; INT4 is the nearest structured label because MQ4V2 has no dedicated enum. Hipfire DFlash2, Q8 KV, runtime block 16, temp 0. All outputs byte-identical and correct merge-sort code; 157 tokens, tau 13.1818, acceptance 0.8788. The 128K Q8-KV point could not load: hipMalloc OOM after target+draft reached 21.32 GB.
- localmaxxing · observed command
- HIP_VISIBLE_DEVICES=0 HIPFIRE_HFQ4G256_LDSSTAGE=1 HIPFIRE_VERIFY_GRAPH=0 dflash_spec_demo --target q38-fixed.mq4v2.mq4 --draft qwen38-27b-dflash2-mq4v2.hfq --prompt-file merge_sort_thinking_off.txt --max 256 --temp 0 --no-chatml --kv-mode q8 --ctx 32768 --no-adaptive-b
- localmaxxing · run id
- cmt62jgz20032lk01uo7i0p1d
- localmaxxing · tokenized · arguments
- dflash_spec_demo, --target, q38-fixed.mq4v2.mq4, --draft, qwen38-27b-dflash2-mq4v2.hfq, --prompt-file, merge_sort_thinking_off.txt, --max, 256, --temp, 0, --no-chatml, --kv-mode, q8, --ctx, 32768, --no-adaptive-b
- localmaxxing · tokenized · environment · HIPFIRE HFQ4G256 LDSSTAGE
- 1
- localmaxxing · tokenized · environment · HIPFIRE VERIFY GRAPH
- 0
- localmaxxing · tokenized · environment · HIP VISIBLE DEVICES
- 0
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmt62jgz20032lk01uo7i0p1d ↗ |