Recipe

ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1

ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
8655 (277ff5fff)
Accelerators
1
Tensor parallel
1
Context tokens
262,144
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/ornith-ai/Ornith-1.0-35B-GGUF
Repository
ornith-ai/Ornith-1.0-35B-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmqw5k2wu039dqr01j2yh9kg5

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. llama-server
  2. --model
  3. ornith-1.0-35b-Q4_K_M.gguf
  4. --ctx-size
  5. 262144
  6. --cache-type-k
  7. q4_0
  8. --cache-type-v
  9. q4_0
  10. --parallel
  11. 1
  12. --batch-size
  13. 1024
  14. --ubatch-size
  15. 256
  16. --flash-attn
  17. on
  18. --n-gpu-layers
  19. 999
  20. --jinja
  21. --reasoning-format
  22. deepseek
  23. --reasoning
  24. auto
  25. --no-cache-prompt
  26. --cache-ram
  27. 0
FlagValue
--modelornith-1.0-35b-Q4_K_M.gguf
--ctx-size262144
--cache-type-kq4_0
--cache-type-vq4_0
--parallel1
--batch-size1024
--ubatch-size256
--flash-attnon
--n-gpu-layers999
--reasoning-formatdeepseek
--reasoningauto
--cache-ram0

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1262,144124.194.4observedornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-3090-24gb
id
ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1
model instance id
ornith-ai-ornith-1-0-35b-gguf--q4-k-m
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
ornith-1-0-35b-gguf-q4-k-m-rtx-3090-24gb-llama-cpp-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
-hf, ornith-ai/Ornith-1.0-35B-GGUF:Q4_K_M, --n-gpu-layers, 999, --host, 0.0.0.0, --port, 8080, -c, 262144
container port
8,080
environment · LLAMA CACHE
/root/.cache/huggingface
host port
8,080
image
ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107f0f0857e259672d1cb85b71c2
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
gemma-4-12b-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
synthesized · template
llama-cpp-server-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
llama.cpp
version
8655 (277ff5fff)

serving

max concurrency
1
max context tokens
262,144
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
RTX 3090
localmaxxing · notes
Ornith-1.0 35B GGUF Q4_K_M on 1x RTX 3090 via llama.cpp CUDA, full 262144 context, q4_0 KV cache, flash-attn, no prompt cache, single streaming request, 5 measured 1000-token runs after 1 warmup, no-thinking template forced with chat_template_kwargs enable_thinking=false. single GPU. Performance-window run with GPU power autotune paused and test GPUs capped at 350W. GPU reached 89C and software thermal slowdown was observed in 22 telemetry samples, so later runs tapered. Host is PCIe 3.0/no NVLink; this row uses one card.
localmaxxing · observed command
llama-server --model ornith-1.0-35b-Q4_K_M.gguf --ctx-size 262144 --cache-type-k q4_0 --cache-type-v q4_0 --parallel 1 --batch-size 1024 --ubatch-size 256 --flash-attn on --n-gpu-layers 999 --jinja --reasoning-format deepseek --reasoning auto --no-cache-prompt --cache-ram 0
localmaxxing · run id
cmqw5k2wu039dqr01j2yh9kg5
localmaxxing · tokenized · arguments
llama-server, --model, ornith-1.0-35b-Q4_K_M.gguf, --ctx-size, 262144, --cache-type-k, q4_0, --cache-type-v, q4_0, --parallel, 1, --batch-size, 1024, --ubatch-size, 256, --flash-attn, on, --n-gpu-layers, 999, --jinja, --reasoning-format, deepseek, --reasoning, auto, --no-cache-prompt, --cache-ram, 0
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmqw5k2wu039dqr01j2yh9kg5