Recipe
nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4
nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4Controller-backed candidate with an exact model revision and retained server-init proof of TP4, MTP5, FP8 KV, and full/piecewise graphs. Client timings survive only as controller summaries, so the index has not promoted or replayed them.
Record
- Status
- candidate
- Source
- 0xsero
- Engine
- vllm
- Engine version
- 0.22.0
- Graph
- full-and-piecewise
- Accelerators
- 4
- Tensor parallel
- 4
- Context tokens
- 261,056
- Max concurrency
- 1
- KV cache tokens
- 376,988
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- no
Hugging Face model card
Identity
https://huggingface.co/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4- Repository
- nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Controller configuration
vllm
No container · candidate · controller
Candidate evidence — not a Run contract
Environment
| Variable | Value |
|---|---|
CUDA_DEVICE_ORDER | PCI_BUS_ID |
CUDA_VISIBLE_DEVICES | GPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853 |
ENABLE_CHUNKED_PREFILL | 1 |
ENABLE_MAMBA_CACHE_STOCHASTIC_ROUNDING | 1 |
ENABLE_MTP | 1 |
ENABLE_PREFIX_CACHING | 1 |
GLOO_SOCKET_IFNAME | lo |
GPU_MEMORY_UTILIZATION | 0.97 |
MAMBA_BACKEND | flashinfer |
MAMBA_CACHE_PHILOX_ROUNDS | 5 |
MAX_MODEL_LEN | 262144 |
MAX_NUM_BATCHED_TOKENS | 1024 |
MAX_NUM_SEQS | 2 |
MOE_BACKEND | auto |
NCCL_CUMEM_HOST_ENABLE | 0 |
NCCL_DEBUG | WARN |
NCCL_IB_DISABLE | 1 |
NCCL_MNNVL_ENABLE | 0 |
NCCL_NVLS_ENABLE | 0 |
NCCL_P2P_DISABLE | 1 |
NCCL_P2P_LEVEL | PIX |
NCCL_SOCKET_IFNAME | lo |
NVIDIA_TF32_OVERRIDE | 1 |
OMP_NUM_THREADS | 8 |
PYTORCH_CUDA_ALLOC_CONF | expandable_segments:True |
RUN_FOREGROUND | 1 |
SAFETENSORS_FAST_GPU | 1 |
TORCH_NCCL_ASYNC_ERROR_HANDLING | 1 |
TORCH_NCCL_BLOCKING_WAIT | 1 |
VLLM_ALLOW_LONG_MAX_MODEL_LEN | 1 |
VLLM_ALLREDUCE_USE_SYMM_MEM | 0 |
VLLM_DISABLE_PYNCCL | 0 |
VLLM_ENGINE_READY_TIMEOUT_S | 1800 |
VLLM_FLASHINFER_ALLREDUCE_BACKEND | trtllm |
VLLM_FLASHINFER_MOE_BACKEND | latency |
VLLM_LOGGING_LEVEL | INFO |
VLLM_SKIP_P2P_CHECK | 0 |
VLLM_USE_FLASHINFER_MOE_FP4 | 1 |
VLLM_USE_FLASHINFER_MOE_FP8 | 1 |
VLLM_WORKER_MULTIPROC_METHOD | spawn |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 564 | — | 84.4 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
| 1 | 32,768 | — | 102.4 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
| 1 | 131,128 | — | 101.5 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
| 1 | 131,128 | — | 76.5 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
| 1 | 261,056 | — | 120.7 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
| 1 | 261,056 | — | 81.7 | — | historical | nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 4
- hardware id
- rtx-pro-6000-blackwell-96gb
- id
- nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4
- model instance id
- nvidia-nvidia-nemotron-3-ultra-550b-a55b-nvfp4--nvfp4-mixed
- recipe source
- 0xsero
- schema version
- local-ai-registry/v1
- speed sweep ids
- nemotron-3-ultra-modelopt-mixed-rtxpro6000-vllm-tp4-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- No
engine
- graph mode
- full-and-piecewise
- name
- vllm
- version
- 0.22.0
serving
- kv cache tokens
- 376,988
- max concurrency
- 1
- max context tokens
- 261,056
- tensor parallel
- 4
Provenance & metadata (3)
facts
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |