Recipe
gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1
gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1Release-blocked Gemma-4-12B-it EXL3 3 bpw candidate on one RTX 2000 Ada. Exact 128K C1/C2/C4, deterministic correctness, and CUDA graphs passed in a pinned base image with the pinned TabbyAPI stack installed, but the declared release image still needs an exact replay.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- tabbyapi
- Engine version
- 0.0.1+e632af41
- Graph
- piecewise
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 131,072
- Max concurrency
- 4
- KV cache tokens
- 557,056
- chat
- yes
- reasoning
- no
- tools
- no
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/turboderp/gemma-4-12B-it-exl3- Repository
- turboderp/gemma-4-12B-it-exl3
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
tabbyapi
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
ghcr.io/theroyallab/tabbyapi:cu13@sha256:ffa8388f310d3c8a2727c66d14a76bb663fdb99a2a49e68a43c3bebb5a3e53f1- Digest
sha256:ffa8388f310d3c8a2727c66d14a76bb663fdb99a2a49e68a43c3bebb5a3e53f1- Port
- 5000
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
main.py--config/app/config.yml
Environment
| Variable | Value |
|---|---|
NVIDIA_VISIBLE_DEVICES | all |
Mounts
| Source | Target |
|---|---|
~/.cache/inference-index/models/gemma-4-12b-it-exl3-3bpw | /workspace/models (read-only) |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 130,560 | 418.9 | 20 | 311,678.8 | candidate | gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep |
| 2 | 130,560 | 210.6 | 16.5 | 620,021.5 | candidate | gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep |
| 4 | 130,560 | 105.7 | 11.5 | 1,234,999.9 | candidate | gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- rtx-2000-ada-16gb
- id
- gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1
- model instance id
- turboderp-gemma-4-12b-it-exl3--3-bpw
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- gemma-4-12b-it-exl3-3bpw-rtx2000ada-tabbyapi-tp1-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- No
- tools
- No
- vision
- No
engine
- graph mode
- piecewise
- name
- tabbyapi
- version
- 0.0.1+e632af41
serving
- kv cache tokens
- 557,056
- max concurrency
- 4
- max context tokens
- 131,072
- tensor parallel
- 1
Provenance & metadata (3)
facts
metadata
- launch config · detail
- The referenced TabbyAPI config for the EXL3 3bpw variant was never captured; the mount has been removed rather than pointing at a nonexistent file.
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |