Recipe

mimo-v2-5-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

mimo-v2-5-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
vLLM 0.21.1rc1.dev85+gd87ee1893.d20260518; image ghcr.io/tonyd2wild/mimo-v2.5-tp2-1m-nvfp4kv@sha256:219043cf57a42499ccc514f4de5b96f9e8e401fee47e435ba96ce06c3171f1c8; nvfp4-kv-diffkv mods copied from tonyd2wild/MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark commit 50c4ff8c119a2398c9b2016c4b50221014ccf770
Accelerators
1
Tensor parallel
1
Context tokens
500,000
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/lukealonso/MiMo-V2.5-NVFP4
Repository
lukealonso/MiMo-V2.5-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmr8n6yos00euqr01kqotgbyj

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. vLLM
  2. MiMo
  3. DFlash
  4. TP=2
  5. container,
  6. max_model_len=500000,
  7. gpu_memory_utilization=0.83,
  8. max_num_seqs=6,
  9. max_num_batched_tokens=4096,
  10. NVFP4
  11. weights,
  12. NVFP4
  13. KV
  14. cache,
  15. triton_attn_diffkv,
  16. DFlash
  17. speculative
  18. decoding
  19. num_speculative_tokens=7,
  20. NCCL_PROTO=LL,
  21. NCCL_MAX_NCHANNELS=2

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1500,0001,538.828.41,586.2observedmimo-v2-5-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
mimo-v2-5-nvfp4-dgx-spark-gb10-128gb-vllm-tp1
model instance id
lukealonso-mimo-v2-5-nvfp4--nvfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
mimo-v2-5-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
--model, lukealonso/MiMo-V2.5-NVFP4, --tensor-parallel-size, 1, --host, 0.0.0.0, --port, 8000, --max-model-len, 500000
container port
8,000
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm
version
vLLM 0.21.1rc1.dev85+gd87ee1893.d20260518; image ghcr.io/tonyd2wild/mimo-v2.5-tp2-1m-nvfp4kv@sha256:219043cf57a42499ccc514f4de5b96f9e8e401fee47e435ba96ce06c3171f1c8; nvfp4-kv-diffkv mods copied from tonyd2wild/MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark commit 50c4ff8c119a2398c9b2016c4b50221014ccf770

serving

max concurrency
1
max context tokens
500,000
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
GB10 Grace Blackwell
localmaxxing · notes
MiMo DFlash C1 LocalMax-style row from llama-benchy 0.3.7, pp=2048 tg=128 depth=0 c=1 runs=5 on two DGX Spark/GB10 nodes. NVFP4 weights with true NVFP4 KV cache, triton_attn_diffkv backend, Xiaomi DFlash draft num_speculative_tokens=7. The DFlash repo lacked the vendored nvfp4-kv-diffkv runtime mod; the missing patch bundle was copied from Tony's earlier MiMo TP2 NVFP4-KV repo before launch. Server logs confirmed kv_cache_dtype=nvfp4, triton_attn_diffkv, inline-dequant kernel enabled, WMMA decode precompile, and successful smoke generation. Hostnames, usernames, LAN IPs, and raw logs omitted.
localmaxxing · observed command
vLLM MiMo DFlash TP=2 container, max_model_len=500000, gpu_memory_utilization=0.83, max_num_seqs=6, max_num_batched_tokens=4096, NVFP4 weights, NVFP4 KV cache, triton_attn_diffkv, DFlash speculative decoding num_speculative_tokens=7, NCCL_PROTO=LL, NCCL_MAX_NCHANNELS=2
localmaxxing · run id
cmr8n6yos00euqr01kqotgbyj
localmaxxing · tokenized · arguments
vLLM, MiMo, DFlash, TP=2, container,, max_model_len=500000,, gpu_memory_utilization=0.83,, max_num_seqs=6,, max_num_batched_tokens=4096,, NVFP4, weights,, NVFP4, KV, cache,, triton_attn_diffkv,, DFlash, speculative, decoding, num_speculative_tokens=7,, NCCL_PROTO=LL,, NCCL_MAX_NCHANNELS=2
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:26:09Z

sources

captured atkindurl
2026-08-30T09:26:09Znormalized-recipewww.localmaxxing.com/en/runs/cmr8n6yos00euqr01kqotgbyj