Recipe

nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4

nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4

Controller-backed candidate with an exact model revision and retained server-init proof of TP4, MTP5, FP8 KV, and full/piecewise graphs. Client timings survive only as controller summaries, so the index has not promoted or replayed them.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
0.22.0
Graph
full-and-piecewise
Accelerators
4
Tensor parallel
4
Context tokens
261,056
Max concurrency
1
KV cache tokens
376,988
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Repository
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Controller configuration

vllm

No container · candidate · controller

Candidate evidence — not a Run contract

Environment

VariableValue
CUDA_DEVICE_ORDERPCI_BUS_ID
CUDA_VISIBLE_DEVICESGPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853
ENABLE_CHUNKED_PREFILL1
ENABLE_MAMBA_CACHE_STOCHASTIC_ROUNDING1
ENABLE_MTP1
ENABLE_PREFIX_CACHING1
GLOO_SOCKET_IFNAMElo
GPU_MEMORY_UTILIZATION0.97
MAMBA_BACKENDflashinfer
MAMBA_CACHE_PHILOX_ROUNDS5
MAX_MODEL_LEN262144
MAX_NUM_BATCHED_TOKENS1024
MAX_NUM_SEQS2
MOE_BACKENDauto
NCCL_CUMEM_HOST_ENABLE0
NCCL_DEBUGWARN
NCCL_IB_DISABLE1
NCCL_MNNVL_ENABLE0
NCCL_NVLS_ENABLE0
NCCL_P2P_DISABLE1
NCCL_P2P_LEVELPIX
NCCL_SOCKET_IFNAMElo
NVIDIA_TF32_OVERRIDE1
OMP_NUM_THREADS8
PYTORCH_CUDA_ALLOC_CONFexpandable_segments:True
RUN_FOREGROUND1
SAFETENSORS_FAST_GPU1
TORCH_NCCL_ASYNC_ERROR_HANDLING1
TORCH_NCCL_BLOCKING_WAIT1
VLLM_ALLOW_LONG_MAX_MODEL_LEN1
VLLM_ALLREDUCE_USE_SYMM_MEM0
VLLM_DISABLE_PYNCCL0
VLLM_ENGINE_READY_TIMEOUT_S1800
VLLM_FLASHINFER_ALLREDUCE_BACKENDtrtllm
VLLM_FLASHINFER_MOE_BACKENDlatency
VLLM_LOGGING_LEVELINFO
VLLM_SKIP_P2P_CHECK0
VLLM_USE_FLASHINFER_MOE_FP41
VLLM_USE_FLASHINFER_MOE_FP81
VLLM_WORKER_MULTIPROC_METHODspawn

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
156484.4historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
132,768102.4historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
1131,128101.5historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
1131,12876.5historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
1261,056120.7historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
1261,05681.7historicalnemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4
model instance id
nvidia-nvidia-nemotron-3-ultra-550b-a55b-nvfp4--nvfp4-mixed
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
full-and-piecewise
name
vllm
version
0.22.0

serving

kv cache tokens
376,988
max concurrency
1
max context tokens
261,056
tensor parallel
4
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry