Recipe

qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1

qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
ROCmFPX Kairic Edge v1.2 (205a3e5f40e5542e2f2eb68e3d3f81f918b1d895)
Accelerators
1
Tensor parallel
1
Context tokens
32,768
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/jcbtc/Qwen3.8-27B-IU4-Kairic-Edge
Repository
jcbtc/Qwen3.8-27B-IU4-Kairic-Edge
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmt5yl3p2000blk01hiumzfjw

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. ./run-kairic-edge-lucebox.sh

Environment

VariableValue
CACHE_RAM8192
CONTEXT32768
PORT18087
ROCR_VISIBLE_DEVICES1

Source notes

llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0

request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
132,768168.590.1observedqwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
ryzen-ai-max-plus-395-128gb
id
qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1
model instance id
jcbtc-qwen3-8-27b-iu4-kairic-edge--iu4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1-sweep
status
candidate

capabilities

engine

name
llama.cpp
version
ROCmFPX Kairic Edge v1.2 (205a3e5f40e5542e2f2eb68e3d3f81f918b1d895)

serving

max concurrency
1
max context tokens
32,768
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-09-01T10:20:39Z
serving.max concurrency · reason explicit-source-server-capacity
serving.max concurrency · state known
serving.max context tokens · provenance · captured at
2026-09-01T10:20:39Z
serving.max context tokens · reason explicit-source-context-limit
serving.max context tokens · state known

metadata

localmaxxing · backend
rocm
localmaxxing · batch size
1
localmaxxing · hardware label
Ryzen AI Max 395
localmaxxing · notes
Forced 512-token natural-prose workload on the Radeon 8060S iGPU (gfx1151) of a BOSGAME M5 Ryzen AI MAX+ 395 with 128GB unified memory. The discrete R9700 was present but idle and not used; ROCR_VISIBLE_DEVICES selected the physical iGPU. Kairic Edge IU4 model plus its published FFN/GDN/GDN-output sidecars, safe strict-compact M65 verifier, n-gram speculation, and MTP4. One warmup discarded, then three measured decode results: 89.9854 / 90.1352 / 92.4784 tok/s; submitted value is the median 90.1352. Total throughput 94.5068 tok/s is the median of each measured request's (prompt + completion tokens) / client wall time. Draft acceptance was 482/534 = 90.26%; all three 512-token outputs were byte-identical and visibly coherent. This result benefits heavily from Kairic's n-gram/speculative path and should not be read as universal model throughput. Per-request prompt caching was disabled. Peak sampled iGPU socket graphics package power during the overall battery was 128.081W, but power is omitted from the structured metric because it was not isolated to this workload. Runtime used a portable ROCm 7.14/GCC 15 build of Kairic v1.2 rather than the author's certified ROCm 7.15/GCC 13 toolchain.
localmaxxing · observed command
ROCR_VISIBLE_DEVICES=1 CONTEXT=32768 CACHE_RAM=8192 PORT=18087 ./run-kairic-edge-lucebox.sh # llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0 # request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs
localmaxxing · run id
cmt5yl3p2000blk01hiumzfjw
localmaxxing · tokenized · arguments
./run-kairic-edge-lucebox.sh
localmaxxing · tokenized · environment · CACHE RAM
8192
localmaxxing · tokenized · environment · CONTEXT
32768
localmaxxing · tokenized · environment · PORT
18087
localmaxxing · tokenized · environment · ROCR VISIBLE DEVICES
1
localmaxxing · tokenized · fidelity
faithful
localmaxxing · tokenized · notes
llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0, request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmt5yl3p2000blk01hiumzfjw