Recipe
qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1
qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- llama.cpp
- Engine version
- ROCmFPX Kairic Edge v1.2 (205a3e5f40e5542e2f2eb68e3d3f81f918b1d895)
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 32,768
- Max concurrency
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/jcbtc/Qwen3.8-27B-IU4-Kairic-Edge- Repository
- jcbtc/Qwen3.8-27B-IU4-Kairic-Edge
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
llama.cpp
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
./run-kairic-edge-lucebox.sh
Environment
| Variable | Value |
|---|---|
CACHE_RAM | 8192 |
CONTEXT | 32768 |
PORT | 18087 |
ROCR_VISIBLE_DEVICES | 1 |
Source notes
llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0
request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 32,768 | 168.5 | 90.1 | — | observed | qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- ryzen-ai-max-plus-395-128gb
- id
- qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1
- model instance id
- jcbtc-qwen3-8-27b-iu4-kairic-edge--iu4
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen3-8-27b-iu4-ryzen-ai-max-plus-395-128gb-llama-cpp-tp1-sweep
- status
- candidate
capabilities
engine
- name
- llama.cpp
- version
- ROCmFPX Kairic Edge v1.2 (205a3e5f40e5542e2f2eb68e3d3f81f918b1d895)
serving
- max concurrency
- 1
- max context tokens
- 32,768
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
- serving.max concurrency · provenance · captured at
- 2026-09-01T10:20:39Z
- serving.max context tokens · provenance · captured at
- 2026-09-01T10:20:39Z
metadata
- localmaxxing · backend
- rocm
- localmaxxing · batch size
- 1
- localmaxxing · hardware label
- Ryzen AI Max 395
- localmaxxing · notes
- Forced 512-token natural-prose workload on the Radeon 8060S iGPU (gfx1151) of a BOSGAME M5 Ryzen AI MAX+ 395 with 128GB unified memory. The discrete R9700 was present but idle and not used; ROCR_VISIBLE_DEVICES selected the physical iGPU. Kairic Edge IU4 model plus its published FFN/GDN/GDN-output sidecars, safe strict-compact M65 verifier, n-gram speculation, and MTP4. One warmup discarded, then three measured decode results: 89.9854 / 90.1352 / 92.4784 tok/s; submitted value is the median 90.1352. Total throughput 94.5068 tok/s is the median of each measured request's (prompt + completion tokens) / client wall time. Draft acceptance was 482/534 = 90.26%; all three 512-token outputs were byte-identical and visibly coherent. This result benefits heavily from Kairic's n-gram/speculative path and should not be read as universal model throughput. Per-request prompt caching was disabled. Peak sampled iGPU socket graphics package power during the overall battery was 128.081W, but power is omitted from the structured metric because it was not isolated to this workload. Runtime used a portable ROCm 7.14/GCC 15 build of Kairic v1.2 rather than the author's certified ROCm 7.15/GCC 13 toolchain.
- localmaxxing · observed command
- ROCR_VISIBLE_DEVICES=1 CONTEXT=32768 CACHE_RAM=8192 PORT=18087 ./run-kairic-edge-lucebox.sh # llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0 # request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs
- localmaxxing · run id
- cmt5yl3p2000blk01hiumzfjw
- localmaxxing · tokenized · arguments
- ./run-kairic-edge-lucebox.sh
- localmaxxing · tokenized · environment · CACHE RAM
- 8192
- localmaxxing · tokenized · environment · CONTEXT
- 32768
- localmaxxing · tokenized · environment · PORT
- 18087
- localmaxxing · tokenized · environment · ROCR VISIBLE DEVICES
- 1
- localmaxxing · tokenized · fidelity
- faithful
- localmaxxing · tokenized · notes
- llama-server: -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -np 1 --kairic-edge --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 --temp 0 --top-p 1 --top-k 0 --min-p 0, request: max_tokens=512, cache_prompt=false; one warmup discarded, then three measured runs
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmt5yl3p2000blk01hiumzfjw ↗ |