Recipe

ornith15-35b-a3b-mixed-iq3-rtx4000ada-llamacpp-tp1

ornith15-35b-a3b-mixed-iq3-rtx4000ada-llamacpp-tp1

Blocked GPU-resident screen for Ornith-1.5-35B-A3B mixed IQ3 on one RTX 4000 Ada. The exact 15,512,189,120-byte GGUF and llama.cpp image are pinned. The prior SGLang attempt never loaded this architecture, so it rejects only that runtime path; this llama.cpp path still needs correctness, exact-128K context, KV capacity, and speed acceptance. CLI launch remains disabled.

Record

Status
candidate
Source
0xsero
Engine
llama-cpp
Engine version
b10481 (25ae3a9b331fffea50ff8d07a5cad34c33f1276f)
Graph
not-applicable
Accelerators
1
Tensor parallel
1
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/AtomicChat/Ornith-1.5-35B-A3B-GGUF
Repository
AtomicChat/Ornith-1.5-35B-A3B-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

llama-cpp

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107ff48e4720145911c86930f2ce
Digest
sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107ff48e4720145911c86930f2ce
Port
8080

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. --model
  2. /models/Ornith-1.5-35B-A3B-AD-IQ3_S-IQ3_XXS.gguf
  3. --host
  4. 0.0.0.0
  5. --port
  6. 8080
  7. --ctx-size
  8. 131072
  9. --parallel
  10. 1
  11. --n-gpu-layers
  12. 99
  13. --flash-attn
  14. on
  15. --cache-type-k
  16. q4_0
  17. --cache-type-v
  18. q4_0
  19. --jinja
FlagValue
--model/models/Ornith-1.5-35B-A3B-AD-IQ3_S-IQ3_XXS.gguf
--host0.0.0.0
--port8080
--ctx-size131072
--parallel1
--n-gpu-layers99
--flash-attnon
--cache-type-kq4_0
--cache-type-vq4_0

Mounts

SourceTarget
~/.cache/inference-index/models/ornith15-mixed-iq3/models (read-only)

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-4000-ada-20gb
id
ornith15-35b-a3b-mixed-iq3-rtx4000ada-llamacpp-tp1
model instance id
atomicchat-ornith-1-5-35b-a3b-gguf--mixed-iq3-s-iq3-xxs
recipe source
0xsero
schema version
local-ai-registry/v1
status
candidate

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
not-applicable
name
llama-cpp
version
b10481 (25ae3a9b331fffea50ff8d07a5cad34c33f1276f)

serving

tensor parallel
1
Provenance & metadata (3)

facts

launch.environment · provenance · captured at
2026-08-27T06:04:13.773Z
launch.environment · reason not-observed
launch.environment · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max concurrency · reason not-observed
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
speed sweep ids · provenance · captured at
2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry