Recipe

qwen38-q4km-arcb70-llamacpp-tp2

qwen38-q4km-arcb70-llamacpp-tp2

Two-B70 Qwen3.8-27B Q4_K_M llama.cpp SYCL candidate. A fresh live screen found four 65,536-token slots, correct chat output, and C1/C2/C4 measurements through 8K, but the checked-in parallel-4 profile cannot expose the claimed 128K per-request context and becomes strongly asymmetric under concurrency.

Record

Status
candidate
Source
0xsero
Engine
llama-cpp
Engine version
4302fb59969a5d8cf9f8e5f55fdd4506d0ed2126+b70-patches
Graph
not-applicable
Accelerators
2
Tensor parallel
2
Context tokens
8,192
Max concurrency
4
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF
Repository
ggml-org/Qwen3.8-27B-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

llama-cpp

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8010

Environment

VariableValue
BATCH8192
DRAFT_NGL0
DRAFT_THREADS16
ENABLE_MTP1
ENABLE_VISION0
GPU_COUNT2
PARALLEL4
SPEC_DRAFT_N_MAX5
THREADS16
UBATCH8192

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
12,50095551.1historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
110,00088150.9historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
120,00080151.1historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
140,00069449.7historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
160,00056346.9historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
1160,00036542historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
1245,00027530.8historicalqwen38-q4km-arcb70-llamacpp-tp2-sweep
11,024406.235.7candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep
21,024358.913.3candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep
41,024201.79candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep
18,192361.927candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep
28,192367.59.1candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep
48,192258.15.4candidateqwen38-q4km-arcb70-llamacpp-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
intel-arc-pro-b70-32gb
id
qwen38-q4km-arcb70-llamacpp-tp2
model instance id
ggml-org-qwen3-8-27b-gguf--q4-k-m
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
qwen38-q4km-arcb70-llamacpp-tp2-sweep
status
candidate

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
not-applicable
name
llama-cpp
version
4302fb59969a5d8cf9f8e5f55fdd4506d0ed2126+b70-patches

serving

max concurrency
4
max context tokens
8,192
tensor parallel
2
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry