Recipe

minimax-m3-mxfp4-rtxpro6000-vllm-tp4

minimax-m3-mxfp4-rtxpro6000-vllm-tp4

Pinned-source MiniMax-M3 MXFP4 TP4 candidate with the required sparse-attention, parser, and clamped-SwiGLU patches. The source reports 250K context and multi-request throughput, but its locally built image and model revision are not immutable, so CLI launch and promotion remain blocked.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
source-built patched image; base commit unreported
Graph
unknown
Accelerators
4
Tensor parallel
4
Context tokens
250,000
Max concurrency
4
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/olka-fi/MiniMax-M3-MXFP4
Repository
olka-fi/MiniMax-M3-MXFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUDA_VISIBLE_DEVICES0,1,2,3
VLLM_ALLOW_LONG_MAX_MODEL_LEN1
VLLM_MXFP4_USE_MARLIN1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1113historicalminimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
446historicalminimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
1131,0722,800historicalminimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
minimax-m3-mxfp4-rtxpro6000-vllm-tp4
model instance id
olka-fi-minimax-m3-mxfp4--mxfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
unknown
name
vllm
version
source-built patched image; base commit unreported

serving

max concurrency
4
max context tokens
250,000
tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry