Recipe

gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1

gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1

Release-blocked Gemma-4-12B-it EXL3 3 bpw candidate on one RTX 2000 Ada. Exact 128K C1/C2/C4, deterministic correctness, and CUDA graphs passed in a pinned base image with the pinned TabbyAPI stack installed, but the declared release image still needs an exact replay.

Record

Status
candidate
Source
0xsero
Engine
tabbyapi
Engine version
0.0.1+e632af41
Graph
piecewise
Accelerators
1
Tensor parallel
1
Context tokens
131,072
Max concurrency
4
KV cache tokens
557,056
chat
yes
reasoning
no
tools
no
vision
no

Hugging Face model card

Identity

https://huggingface.co/turboderp/gemma-4-12B-it-exl3
Repository
turboderp/gemma-4-12B-it-exl3
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

tabbyapi

Container · candidate · docker

Candidate evidence — not a Run contract

Image
ghcr.io/theroyallab/tabbyapi:cu13@sha256:ffa8388f310d3c8a2727c66d14a76bb663fdb99a2a49e68a43c3bebb5a3e53f1
Digest
sha256:ffa8388f310d3c8a2727c66d14a76bb663fdb99a2a49e68a43c3bebb5a3e53f1
Port
5000

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. main.py
  2. --config
  3. /app/config.yml

Environment

VariableValue
NVIDIA_VISIBLE_DEVICESall

Mounts

SourceTarget
~/.cache/inference-index/models/gemma-4-12b-it-exl3-3bpw/workspace/models (read-only)

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1130,560418.920311,678.8candidategemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep
2130,560210.616.5620,021.5candidategemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep
4130,560105.711.51,234,999.9candidategemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
rtx-2000-ada-16gb
id
gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1
model instance id
turboderp-gemma-4-12b-it-exl3--3-bpw
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
No
tools
No
vision
No

engine

graph mode
piecewise
name
tabbyapi
version
0.0.1+e632af41

serving

kv cache tokens
557,056
max concurrency
4
max context tokens
131,072
tensor parallel
1
Provenance & metadata (3)

facts

metadata

launch config · detail
The referenced TabbyAPI config for the EXL3 3bpw variant was never captured; the mount has been removed rather than pointing at a nonexistent file.
launch config · state not-captured

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry