Recipe

nvidia-nemotron-labs-3-puzzle-75b-a9b-bf16-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

nvidia-nemotron-labs-3-puzzle-75b-a9b-bf16-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
0.24.0+aeon.sm121a.dflash
Accelerators
1
Tensor parallel
1
Context tokens
262,144
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4
Repository
nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmrbhpcxc00krro01oq0feln3

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. docker
  2. run
  3. -d
  4. --name
  5. nemotron-labs3-vllm
  6. --gpus
  7. all
  8. --ipc
  9. host
  10. --ulimit
  11. memlock=-1
  12. --ulimit
  13. stack=67108864
  14. --security-opt
  15. seccomp=unconfined
  16. -p
  17. 8000:8000
  18. -v
  19. /home/neo/models/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4:/model:ro
  20. -e
  21. NVIDIA_VISIBLE_DEVICES=all
  22. -e
  23. NVIDIA_DRIVER_CAPABILITIES=compute,utility
  24. --entrypoint
  25. /usr/local/bin/python3.12
  26. ghcr.io/aeon-7/aeon-vllm-ultimate:latest
  27. /usr/local/bin/vllm
  28. serve
  29. /model
  30. --served-model-name
  31. nemotron-labs3-puzzle-75b-a9b-nvfp4
  32. --port
  33. 8000
  34. --tensor-parallel-size
  35. 1
  36. --enable-expert-parallel
  37. --async-scheduling
  38. --trust-remote-code
  39. --mamba-backend
  40. flashinfer
  41. --mamba_ssm_cache_dtype
  42. float16
  43. --enable-mamba-cache-stochastic-rounding
  44. --mamba-cache-philox-rounds
  45. 5
  46. --speculative-config
  47. {"method":"mtp","num_speculative_tokens":3}
  48. --tool-call-parser
  49. qwen3_coder
  50. --reasoning-parser
  51. nemotron_v3
  52. --enable-auto-tool-choice
FlagValue
--namenemotron-labs3-vllm
--gpusall
--ipchost
--ulimitmemlock=-1
--ulimitstack=67108864
--security-optseccomp=unconfined
-p8000:8000
-v/home/neo/models/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4:/model:ro
-eNVIDIA_VISIBLE_DEVICES=all
-eNVIDIA_DRIVER_CAPABILITIES=compute,utility
--entrypoint/usr/local/bin/python3.12
--served-model-namenemotron-labs3-puzzle-75b-a9b-nvfp4
--port8000
--tensor-parallel-size1
--mamba-backendflashinfer
--mamba_ssm_cache_dtypefloat16
--mamba-cache-philox-rounds5
--speculative-config{"method":"mtp","num_speculative_tokens":3}
--tool-call-parserqwen3_coder
--reasoning-parsernemotron_v3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
262,144231.6observednvidia-nemotron-labs-3-puzzle-75b-a9b-bf16-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
nvidia-nemotron-labs-3-puzzle-75b-a9b-bf16-nvfp4-dgx-spark-gb10-128gb-vllm-tp1
model instance id
nvidia-nvidia-nemotron-labs-3-puzzle-75b-a9b-nvfp4--nvfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
nvidia-nemotron-labs-3-puzzle-75b-a9b-bf16-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
serve, --model, nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4, --tensor-parallel-size, 1, --host, 0.0.0.0, --port, 8000, --max-model-len, 262144
container port
8,000
entrypoint
vllm
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm
version
0.24.0+aeon.sm121a.dflash

serving

max context tokens
262,144
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
draft launch.entrypoint · provenance · captured at
2026-09-01T01:41:26Z
draft launch.entrypoint · reason linux-amd64-container-config-entrypoint
draft launch.entrypoint · state known
engine.graph mode · provenance · captured at
2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-not-evidenced
serving.max concurrency · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · batch size
8
localmaxxing · hardware label
GB10 Grace Blackwell
localmaxxing · notes
DGX Spark / NVIDIA GB10 local vLLM benchmark. Nemotron Labs 3 Puzzle 75B A9B NVFP4 served from local HF snapshot with MTP speculative decoding enabled; greedy / temperature 0; warmed server; selected best successful run by output tok/s from 512:128, 4096:128, 512:512 x concurrency 1,2,4,8 sweep. Linux reports ~121.7 GiB visible memory; platform class is 128 GB unified memory.
localmaxxing · observed command
docker run -d --name nemotron-labs3-vllm --gpus all --ipc host --ulimit memlock=-1 --ulimit stack=67108864 --security-opt seccomp=unconfined -p 8000:8000 -v /home/neo/models/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4:/model:ro -e NVIDIA_VISIBLE_DEVICES=all -e NVIDIA_DRIVER_CAPABILITIES=compute,utility --entrypoint /usr/local/bin/python3.12 ghcr.io/aeon-7/aeon-vllm-ultimate:latest /usr/local/bin/vllm serve /model --served-model-name nemotron-labs3-puzzle-75b-a9b-nvfp4 --port 8000 --tensor-parallel-size 1 --enable-expert-parallel --async-scheduling --trust-remote-code --mamba-backend flashinfer --mamba_ssm_cache_dtype float16 --enable-mamba-cache-stochastic-rounding --mamba-cache-philox-rounds 5 --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --tool-call-parser qwen3_coder --reasoning-parser nemotron_v3 --enable-auto-tool-choice
localmaxxing · run id
cmrbhpcxc00krro01oq0feln3
localmaxxing · tokenized · arguments
docker, run, -d, --name, nemotron-labs3-vllm, --gpus, all, --ipc, host, --ulimit, memlock=-1, --ulimit, stack=67108864, --security-opt, seccomp=unconfined, -p, 8000:8000, -v, /home/neo/models/nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4:/model:ro, -e, NVIDIA_VISIBLE_DEVICES=all, -e, NVIDIA_DRIVER_CAPABILITIES=compute,utility, --entrypoint, /usr/local/bin/python3.12, ghcr.io/aeon-7/aeon-vllm-ultimate:latest, /usr/local/bin/vllm, serve, /model, --served-model-name, nemotron-labs3-puzzle-75b-a9b-nvfp4, --port, 8000, --tensor-parallel-size, 1, --enable-expert-parallel, --async-scheduling, --trust-remote-code, --mamba-backend, flashinfer, --mamba_ssm_cache_dtype, float16, --enable-mamba-cache-stochastic-rounding, --mamba-cache-philox-rounds, 5, --speculative-config, {"method":"mtp","num_speculative_tokens":3}, --tool-call-parser, qwen3_coder, --reasoning-parser, nemotron_v3, --enable-auto-tool-choice
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:26:09Z

sources

captured atkindurl
2026-08-30T09:26:09Znormalized-recipewww.localmaxxing.com/en/runs/cmrbhpcxc00krro01oq0feln3