Recipe
ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1
ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- llama.cpp
- Engine version
- 8655 (277ff5fff)
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 262,144
- Max concurrency
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/ornith-ai/Ornith-1.0-35B-GGUF- Repository
- ornith-ai/Ornith-1.0-35B-GGUF
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
llama.cpp
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
llama-server--modelornith-1.0-35b-Q4_K_M.gguf--ctx-size262144--cache-type-kq4_0--cache-type-vq4_0--parallel1--batch-size1024--ubatch-size256--flash-attnon--n-gpu-layers999--jinja--reasoning-formatdeepseek--reasoningauto--no-cache-prompt--cache-ram0
| Flag | Value |
|---|---|
--model | ornith-1.0-35b-Q4_K_M.gguf |
--ctx-size | 262144 |
--cache-type-k | q4_0 |
--cache-type-v | q4_0 |
--parallel | 1 |
--batch-size | 1024 |
--ubatch-size | 256 |
--flash-attn | on |
--n-gpu-layers | 999 |
--reasoning-format | deepseek |
--reasoning | auto |
--cache-ram | 0 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 262,144 | — | 124.1 | 94.4 | observed | ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- rtx-3090-24gb
- id
- ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1
- model instance id
- ornith-ai-ornith-1-0-35b-gguf--q4-k-m
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- -hf, ornith-ai/Ornith-1.0-35B-GGUF:Q4_K_M, --n-gpu-layers, 999, --host, 0.0.0.0, --port, 8080, -c, 262144
- container port
- 8,080
- environment · LLAMA CACHE
- /root/.cache/huggingface
- host port
- 8,080
- image
- ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107f0f0857e259672d1cb85b71c2
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- gemma-4-12b-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
- synthesized · template
- llama-cpp-server-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- llama.cpp
- version
- 8655 (277ff5fff)
serving
- max concurrency
- 1
- max context tokens
- 262,144
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
metadata
- localmaxxing · backend
- cuda
- localmaxxing · hardware label
- RTX 3090
- localmaxxing · notes
- Ornith-1.0 35B GGUF Q4_K_M on 1x RTX 3090 via llama.cpp CUDA, full 262144 context, q4_0 KV cache, flash-attn, no prompt cache, single streaming request, 5 measured 1000-token runs after 1 warmup, no-thinking template forced with chat_template_kwargs enable_thinking=false. single GPU. Performance-window run with GPU power autotune paused and test GPUs capped at 350W. GPU reached 89C and software thermal slowdown was observed in 22 telemetry samples, so later runs tapered. Host is PCIe 3.0/no NVLink; this row uses one card.
- localmaxxing · observed command
- llama-server --model ornith-1.0-35b-Q4_K_M.gguf --ctx-size 262144 --cache-type-k q4_0 --cache-type-v q4_0 --parallel 1 --batch-size 1024 --ubatch-size 256 --flash-attn on --n-gpu-layers 999 --jinja --reasoning-format deepseek --reasoning auto --no-cache-prompt --cache-ram 0
- localmaxxing · run id
- cmqw5k2wu039dqr01j2yh9kg5
- localmaxxing · tokenized · arguments
- llama-server, --model, ornith-1.0-35b-Q4_K_M.gguf, --ctx-size, 262144, --cache-type-k, q4_0, --cache-type-v, q4_0, --parallel, 1, --batch-size, 1024, --ubatch-size, 256, --flash-attn, on, --n-gpu-layers, 999, --jinja, --reasoning-format, deepseek, --reasoning, auto, --no-cache-prompt, --cache-ram, 0
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:10:02Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:10:02Z | normalized-recipe | www.localmaxxing.com/en/runs/cmqw5k2wu039dqr01j2yh9kg5 ↗ |