Recipe

nex-n2-pro-nvfp4-rtxpro6000-sglang-tp4

nex-n2-pro-nvfp4-rtxpro6000-sglang-tp4

Controller-backed NEXTN candidate. Exact baseline model/image pins and a 144K raw baseline sweep were recovered, but the later NEXTN graft and its 81-94 tok/s summaries are not immutably pinned; baseline evidence does not validate this leaf.

Record

Status
candidate
Source
0xsero
Engine
sglang
Engine version
0.5.12 fork
Graph
unknown
Accelerators
4
Tensor parallel
4
Context tokens
103,000
Max concurrency
1
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/local/Nex-N2-Pro-NVFP4
Repository
local/Nex-N2-Pro-NVFP4
Status
unknown
Link type
Exact Hub repository

hf api access unresolved

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Controller configuration

sglang

No container · candidate · controller

Candidate evidence — not a Run contract

Environment

VariableValue
CUDA_DEVICE_ORDERPCI_BUS_ID
CUDA_VISIBLE_DEVICESGPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853
SPEC_ALGONEXTN
SPEC_NUM_DRAFT_TOKENS4

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
172.6historicalnex-n2-pro-nvfp4-rtxpro6000-sglang-tp4-sweep
181.1historicalnex-n2-pro-nvfp4-rtxpro6000-sglang-tp4-sweep
193.6historicalnex-n2-pro-nvfp4-rtxpro6000-sglang-tp4-sweep
189.2historicalnex-n2-pro-nvfp4-rtxpro6000-sglang-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
nex-n2-pro-nvfp4-rtxpro6000-sglang-tp4
model instance id
local-nex-n2-pro-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
nex-n2-pro-nvfp4-rtxpro6000-sglang-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
unknown
name
sglang
version
0.5.12 fork

serving

max concurrency
1
max context tokens
103,000
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry