Speed sweep

minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep

minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep

Record

Recipe
minimax-m3-mxfp4-rtxpro6000-vllm-tp4
Measured
2026-08-24
Accepted
2026-08-24
Points
3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
1113historical
446historical
1131,0722,800historical

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

accepted at
2026-08-24
id
minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
measured at
2026-08-24
recipe id
minimax-m3-mxfp4-rtxpro6000-vllm-tp4
schema version
local-ai-registry/v1

metrics

concurrency
4
inference engine version
source-built patched image; base commit unreported
latest point at
2026-08-24
max context tokens
131,072
peak generation tps
184
peak prompt tps
2,800
point count
3

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
1Unknown113113UnknownUnknownUnknown1
4Unknown18446UnknownUnknownUnknown1
1131,072UnknownUnknownUnknownUnknown2,8001
Provenance & metadata (1)

source

commit
303701cd177e6adecd4999beb67f4797a24439db
paths
README.md, Dockerfile, docker-compose.yml, scripts/serve.sh, patches/vllm/models/minimax_m3/nvidia/model.py, patches/vllm/model_executor/layers/quantization/compressed_tensors/compressed_tensors_moe/compressed_tensors_moe_w4a4_mxfp4.py