Recipe

ornith15-35b-a3b-bf16-mi300x-sglang-tp1

ornith15-35b-a3b-bf16-mi300x-sglang-tp1

Ornith-1.5-35B-A3B BF16 on one AMD Instinct MI300X with exact 128K C1/C2/C4 acceptance

Record

Status
validated
Source
0xsero
Engine
sglang
Engine version
0.5.18.dev20260823+gdd15fb57b5
Graph
full
Accelerators
1
Tensor parallel
1
Context tokens
131,072
Max concurrency
4
KV cache tokens
2,226,502
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B
Repository
ornith-ai/Ornith-1.5-35B-A3B
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Validated: pinned artifact, pinned runtime, and accepted evidence. This is a launch contract.

Docker configuration

sglang

Container · validated · docker

Validated launch contract

Image
lmsysorg/sglang-rocm:v0.5.18-rocm720-mi30x-20260823@sha256:8805fe3744f9b549b4e329b37134bce6ce4afd7ec749ea1249507bcbcea10096
Digest
sha256:8805fe3744f9b549b4e329b37134bce6ce4afd7ec749ea1249507bcbcea10096
Port
30000

Launch arguments

  1. -m
  2. sglang.launch_server
  3. --model-path
  4. ornith-ai/Ornith-1.5-35B-A3B
  5. --revision
  6. 10fbf86fed7ecee4a061f8b499a618f46001cac1
  7. --served-model-name
  8. ornith-1.5-35b-a3b
  9. --trust-remote-code
  10. --host
  11. 0.0.0.0
  12. --port
  13. 30000
  14. --context-length
  15. 262144
  16. --mem-fraction-static
  17. 0.90
  18. --max-running-requests
  19. 4
  20. --attention-backend
  21. aiter
  22. --mamba-ssm-dtype
  23. bfloat16
  24. --mamba-radix-cache-strategy
  25. extra_buffer
  26. --chunked-prefill-size
  27. 2048
  28. --max-prefill-tokens
  29. 16384
  30. --cuda-graph-max-bs-decode
  31. 4
  32. --enable-cache-report
  33. --reasoning-parser
  34. qwen3
  35. --tool-call-parser
  36. qwen3_coder
FlagValue
-msglang.launch_server
--model-pathornith-ai/Ornith-1.5-35B-A3B
--revision10fbf86fed7ecee4a061f8b499a618f46001cac1
--served-model-nameornith-1.5-35b-a3b
--host0.0.0.0
--port30000
--context-length262144
--mem-fraction-static0.90
--max-running-requests4
--attention-backendaiter
--mamba-ssm-dtypebfloat16
--mamba-radix-cache-strategyextra_buffer
--chunked-prefill-size2048
--max-prefill-tokens16384
--cuda-graph-max-bs-decode4
--reasoning-parserqwen3
--tool-call-parserqwen3_coder

Environment

VariableValue
HF_HOME/root/.cache/huggingface

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Launch

Exact materialization of this validated launch contract: digest-pinned image, pinned model revision, and the audited arguments. Self-contained — required assets are fetched from this registry and verified against their recorded sha256 before mounting. Also available as local-ai run ornith15-35b-a3b-bf16-mi300x-sglang-tp1.

  1. Pulls the exact container image by sha256 digest — the bytes that were validated, not a floating tag.
  2. Fetches any required engine assets from this registry and verifies each against its recorded sha256; the audited launch script verifies them again inside the container before use.
  3. Downloads the pinned model revision into your Hugging Face cache on first run (reused afterwards).
  4. Serves an OpenAI-compatible API on localhost:30000 — point any client at it.
docker run --rm \
  --ipc host \
  --shm-size 16g \
  -p 30000:30000 \
  -e HF_HOME=/root/.cache/huggingface \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  --entrypoint /opt/venv/bin/python \
  lmsysorg/sglang-rocm:v0.5.18-rocm720-mi30x-20260823@sha256:8805fe3744f9b549b4e329b37134bce6ce4afd7ec749ea1249507bcbcea10096 \
  -m \
  sglang.launch_server \
  --model-path \
  ornith-ai/Ornith-1.5-35B-A3B \
  --revision \
  10fbf86fed7ecee4a061f8b499a618f46001cac1 \
  --served-model-name \
  ornith-1.5-35b-a3b \
  --trust-remote-code \
  --host \
  0.0.0.0 \
  --port \
  30000 \
  --context-length \
  262144 \
  --mem-fraction-static \
  0.90 \
  --max-running-requests \
  4 \
  --attention-backend \
  aiter \
  --mamba-ssm-dtype \
  bfloat16 \
  --mamba-radix-cache-strategy \
  extra_buffer \
  --chunked-prefill-size \
  2048 \
  --max-prefill-tokens \
  16384 \
  --cuda-graph-max-bs-decode \
  4 \
  --enable-cache-report \
  --reasoning-parser \
  qwen3 \
  --tool-call-parser \
  qwen3_coder

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1130,56012,225.596.810,729.9acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
1130,560636,78096.8200.1acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
2130,5604616,230.3acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
2130,56091.2342.6acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
4130,56030.231,675.1acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
4130,56083.1385.2acceptedornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
mi300x-192gb
id
ornith15-35b-a3b-bf16-mi300x-sglang-tp1
model instance id
ornith-ai-ornith-1-5-35b-a3b--bf16
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
ornith15-35b-a3b-bf16-mi300x-sglang-tp1-sweep
status
validated

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
full
name
sglang
version
0.5.18.dev20260823+gdd15fb57b5

serving

kv cache tokens
2,226,502
max concurrency
4
max context tokens
131,072
tensor parallel
1
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry