Recipe

gemma-4-12b-it-nvfp4-rtx5090-sglang-tp1

gemma-4-12b-it-nvfp4-rtx5090-sglang-tp1

Capacity-limited Gemma-4-12B-it NVFP4 candidate on one RTX 5090 with the checksum-verified Gemma engine patch, exact 128K C1 acceptance, full decode CUDA graphs, and C2/C4 rejected by the measured 133368-token KV envelope

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
0.0.0.dev1+geec794bce
Graph
full
Accelerators
1
Tensor parallel
1
Context tokens
131,072
Max concurrency
1
KV cache tokens
133,368
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/unsloth/gemma-4-12b-it-NVFP4
Repository
unsloth/gemma-4-12b-it-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
lmsysorg/sglang:dev-cu13@sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
Digest
sha256:6cd4635214f279e0a43019f88e3120d407567640a58aa7dcc0085e3d91402cc4
Port
30000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. -lc
  2. set -euo pipefail; echo 'dede8848dbcbfcfc6507da963a65d120add93101b6ff5fb54a03f19e86233d11 /tmp/gemma4-engine-patch.diff' | sha256sum -c -; echo 'd5498b253f35e83ab0aaf219a2c2bf2f42c6bbd3f95ffce02760e78b2e38a4e9 /sgl-workspace/sglang/python/sglang/srt/models/gemma4_unified.py' | sha256sum -c -; (cd /sgl-workspace/sglang && patch -p0 < /tmp/gemma4-engine-patch.diff); echo 'b0614b99a0d7fe654ed102fb5db04c578e3be6042896f73f9418965e6672c737 /sgl-workspace/sglang/python/sglang/srt/models/gemma4_unified.py' | sha256sum -c -; exec /opt/sglang/bin/python -m sglang.launch_server --model-path unsloth/gemma-4-12b-it-NVFP4 --revision b1f649734b34aa5575b03d186abd1b9be3d0d5c4 --tp 1 --host 0.0.0.0 --port 30000 --context-length 131072 --mem-fraction-static 0.8827875 --attention-backend triton --enable-cache-report --trust-remote-code --reasoning-parser gemma4 --tool-call-parser gemma4

Environment

VariableValue
HF_HOME/root/.cache/huggingface

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface
asset/gemma4-nvfp4-sglang-engine-patch.diff/tmp/gemma4-engine-patch.diff (read-only)

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1130,5602,936.839.944,472.5candidategemma-4-12b-it-nvfp4-rtx5090-sglang-tp1-sweep
1130,560532,526.539.9244.6candidategemma-4-12b-it-nvfp4-rtx5090-sglang-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-5090-32gb
id
gemma-4-12b-it-nvfp4-rtx5090-sglang-tp1
model instance id
unsloth-gemma-4-12b-it-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
gemma-4-12b-it-nvfp4-rtx5090-sglang-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
full
name
sglang
version
0.0.0.dev1+geec794bce

serving

kv cache tokens
133,368
max concurrency
1
max context tokens
131,072
tensor parallel
1
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry