Recipe

qwen3-8-27b-q4-k-m-rtx-3090-24gb-ollama-tp1

qwen3-8-27b-q4-k-m-rtx-3090-24gb-ollama-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
ollama
Engine version
0.30.8
Accelerators
1
Tensor parallel
1
Context tokens
32,768
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
Repository
unsloth/Qwen3.8-27B-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

ollama

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmsurfonh06t1ms01sq2vmpx4

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. ollama
  2. serve

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
132,76832540.867.7observedqwen3-8-27b-q4-k-m-rtx-3090-24gb-ollama-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-3090-24gb
id
qwen3-8-27b-q4-k-m-rtx-3090-24gb-ollama-tp1
model instance id
unsloth-qwen3-8-27b-gguf--q4-k-m
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
qwen3-8-27b-q4-k-m-rtx-3090-24gb-ollama-tp1-sweep
status
candidate

capabilities

engine

name
ollama
version
0.30.8

serving

max context tokens
32,768
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · batch size
1
localmaxxing · hardware label
RTX 3090
localmaxxing · notes
Official Ollama library model qwen3.8:27b-q4_k_m. One warmup plus median of three timed 256-token greedy runs. GPU power is the mean of 98 sustained samples at >=90% utilization and >=300 W; peak VRAM sampled every 250 ms.
localmaxxing · observed command
ollama serve
localmaxxing · run id
cmsurfonh06t1ms01sq2vmpx4
localmaxxing · tokenized · arguments
ollama, serve
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmsurfonh06t1ms01sq2vmpx4