Recipe

qwen38-awq-int4-rtx3090-vllm-tp4

qwen38-awq-int4-rtx3090-vllm-tp4

Qwen3.8-27B AWQ INT4 on four RTX 3090s

Record

Status
validated
Source
0xsero
Engine
vllm
Engine version
0.27.1
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
131,072
Max concurrency
4
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/cyankiwi/Qwen3.8-27B-AWQ-INT4
Repository
cyankiwi/Qwen3.8-27B-AWQ-INT4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Validated: pinned artifact, pinned runtime, and accepted evidence. This is a launch contract.

Docker configuration

vllm

Container · validated · docker

Validated launch contract

Image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
Digest
sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
Port
12434

Launch arguments

  1. --model
  2. cyankiwi/Qwen3.8-27B-AWQ-INT4
  3. --revision
  4. 63768c10df38c0395e12ef49edac1bd539eaeeea
  5. --served-model-name
  6. Qwen3.8-27B
  7. --tensor-parallel-size
  8. 4
  9. --max-model-len
  10. 131072
  11. --max-num-seqs
  12. 16
  13. --kv-cache-dtype
  14. fp8
  15. --gpu-memory-utilization
  16. 0.94
  17. --enable-auto-tool-choice
  18. --tool-call-parser
  19. qwen3_xml
  20. --reasoning-parser
  21. qwen3
FlagValue
--modelcyankiwi/Qwen3.8-27B-AWQ-INT4
--revision63768c10df38c0395e12ef49edac1bd539eaeeea
--served-model-nameQwen3.8-27B
--tensor-parallel-size4
--max-model-len131072
--max-num-seqs16
--kv-cache-dtypefp8
--gpu-memory-utilization0.94
--tool-call-parserqwen3_xml
--reasoning-parserqwen3

Environment

VariableValue
VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS0

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface
~/.cache/inference-index/vllm/root/.cache/vllm

Launch

Exact materialization of this validated launch contract: digest-pinned image, pinned model revision, and the audited arguments. Self-contained — required assets are fetched from this registry and verified against their recorded sha256 before mounting. Also available as local-ai run qwen38-awq-int4-rtx3090-vllm-tp4.

  1. Pulls the exact container image by sha256 digest — the bytes that were validated, not a floating tag.
  2. Fetches any required engine assets from this registry and verifies each against its recorded sha256; the audited launch script verifies them again inside the container before use.
  3. Downloads the pinned model revision into your Hugging Face cache on first run (reused afterwards).
  4. Serves an OpenAI-compatible API on localhost:12434 — point any client at it.
docker run --rm \
  --gpus all \
  --ipc host \
  --shm-size 32g \
  -p 12434:8000 \
  -e VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v ~/.cache/inference-index/vllm:/root/.cache/vllm \
  vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967 \
  --model \
  cyankiwi/Qwen3.8-27B-AWQ-INT4 \
  --revision \
  63768c10df38c0395e12ef49edac1bd539eaeeea \
  --served-model-name \
  Qwen3.8-27B \
  --tensor-parallel-size \
  4 \
  --max-model-len \
  131072 \
  --max-num-seqs \
  16 \
  --kv-cache-dtype \
  fp8 \
  --gpu-memory-utilization \
  0.94 \
  --enable-auto-tool-choice \
  --tool-call-parser \
  qwen3_xml \
  --reasoning-parser \
  qwen3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
163.6historicalqwen38-awq-int4-rtx3090-vllm-tp4-sweep
4214.9historicalqwen38-awq-int4-rtx3090-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-3090-24gb
id
qwen38-awq-int4-rtx3090-vllm-tp4
model instance id
cyankiwi-qwen3-8-27b-awq-int4--int4-w4a16
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
qwen38-awq-int4-rtx3090-vllm-tp4-sweep
status
validated

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
full-and-piecewise
name
vllm
version
0.27.1

serving

max concurrency
4
max context tokens
131,072
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry