Recipe
ornith15-35b-a3b-iq2-xxs-rtx2000ada-llamacpp-tp1
ornith15-35b-a3b-iq2-xxs-rtx2000ada-llamacpp-tp1Blocked GPU-resident screen for Ornith-1.5-35B-A3B IQ2_XXS on one RTX 2000 Ada. The exact 10,255,140,512-byte GGUF and llama.cpp image are pinned, but no model load, correctness, exact-128K context, KV capacity, or speed evidence exists on this card. CLI launch remains disabled until a funded acceptance window; CPU offload must be recorded as a separate recipe if needed.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- llama-cpp
- Engine version
- b10481 (25ae3a9b331fffea50ff8d07a5cad34c33f1276f)
- Graph
- not-applicable
- Accelerators
- 1
- Tensor parallel
- 1
- chat
- yes
- reasoning
- no
- tools
- no
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/bartowski/Ornith-1.5-35B-A3B-GGUF- Repository
- bartowski/Ornith-1.5-35B-A3B-GGUF
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
llama-cpp
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107ff48e4720145911c86930f2ce- Digest
sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107ff48e4720145911c86930f2ce- Port
- 8080
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
--model/models/Ornith-1.5-35B-A3B-IQ2_XXS.gguf--host0.0.0.0--port8080--ctx-size131072--parallel1--n-gpu-layers99--flash-attnon--cache-type-kq4_0--cache-type-vq4_0--jinja
| Flag | Value |
|---|---|
--model | /models/Ornith-1.5-35B-A3B-IQ2_XXS.gguf |
--host | 0.0.0.0 |
--port | 8080 |
--ctx-size | 131072 |
--parallel | 1 |
--n-gpu-layers | 99 |
--flash-attn | on |
--cache-type-k | q4_0 |
--cache-type-v | q4_0 |
Mounts
| Source | Target |
|---|---|
~/.cache/inference-index/models/ornith15-iq2-xxs | /models (read-only) |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- rtx-2000-ada-16gb
- id
- ornith15-35b-a3b-iq2-xxs-rtx2000ada-llamacpp-tp1
- model instance id
- bartowski-ornith-1-5-35b-a3b-gguf--iq2-xxs
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- No
- tools
- No
- vision
- No
engine
- graph mode
- not-applicable
- name
- llama-cpp
- version
- b10481 (25ae3a9b331fffea50ff8d07a5cad34c33f1276f)
serving
- tensor parallel
- 1
Provenance & metadata (3)
facts
- launch.environment · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.kv cache tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max concurrency · provenance · captured at
- 2026-08-27T06:04:13.773Z
- serving.max context tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
- speed sweep ids · provenance · captured at
- 2026-08-27T06:04:13.773Z
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |