Recipe

ornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1

ornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1

Capacity-limited Ornith-1.5-35B-A3B NVFP4 candidate on one RTX 5090 with exact 128K C1/C2 speed measurements and full CUDA graphs. C4 exceeds the measured KV envelope, and CLI launch remains blocked until a deterministic generation-health replay is captured.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
0.0.0.dev1+geec794bce
Graph
full-and-piecewise
Accelerators
1
Tensor parallel
1
Context tokens
131,072
Max concurrency
2
KV cache tokens
381,008
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-NVFP4
Repository
ornith-ai/Ornith-1.5-35B-A3B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
Digest
sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
Port
30000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. -m
  2. sglang.launch_server
  3. --model-path
  4. ornith-ai/Ornith-1.5-35B-A3B-NVFP4
  5. --revision
  6. 0f0b1b59b879ccde1353e6ebd0fb10c204d4c544
  7. --tp
  8. 1
  9. --host
  10. 0.0.0.0
  11. --port
  12. 30000
  13. --context-length
  14. 131072
  15. --mem-fraction-static
  16. 0.90
  17. --max-running-requests
  18. 4
  19. --cuda-graph-max-bs-decode
  20. 4
  21. --enable-cache-report
  22. --trust-remote-code
  23. --reasoning-parser
  24. qwen3
  25. --tool-call-parser
  26. qwen3_coder
  27. --kv-cache-dtype
  28. fp8_e4m3
  29. --attention-backend
  30. flashinfer
  31. --moe-runner-backend
  32. flashinfer_cutlass
FlagValue
-msglang.launch_server
--model-pathornith-ai/Ornith-1.5-35B-A3B-NVFP4
--revision0f0b1b59b879ccde1353e6ebd0fb10c204d4c544
--tp1
--host0.0.0.0
--port30000
--context-length131072
--mem-fraction-static0.90
--max-running-requests4
--cuda-graph-max-bs-decode4
--reasoning-parserqwen3
--tool-call-parserqwen3_coder
--kv-cache-dtypefp8_e4m3
--attention-backendflashinfer
--moe-runner-backendflashinfer_cutlass

Environment

VariableValue
HF_HOME/root/.cache/huggingface

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1130,5608,986.4207.314,584.7candidateornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1-sweep
1130,5601,495,084.3207.368.4candidateornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1-sweep
2130,56048.121,792candidateornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1-sweep
2130,560172.7112.7candidateornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-5090-32gb
id
ornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1
model instance id
ornith-ai-ornith-1-5-35b-a3b-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
ornith15-35b-a3b-nvfp4-rtx5090-sglang-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
full-and-piecewise
name
sglang
version
0.0.0.dev1+geec794bce

serving

kv cache tokens
381,008
max concurrency
2
max context tokens
131,072
tensor parallel
1
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry