Speed sweep
minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweepRecord
- Measured
- 2026-08-24
- Accepted
- 2026-08-24
- Points
- 3
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| 1 | — | — | 113 | — | historical |
| 4 | — | — | 46 | — | historical |
| 1 | 131,072 | 2,800 | — | — | historical |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- accepted at
- 2026-08-24
- id
- minimax-m3-mxfp4-rtxpro6000-vllm-tp4-sweep
- measured at
- 2026-08-24
- recipe id
- minimax-m3-mxfp4-rtxpro6000-vllm-tp4
- schema version
- local-ai-registry/v1
metrics
- concurrency
- 4
- inference engine version
- source-built patched image; base commit unreported
- latest point at
- 2026-08-24
- max context tokens
- 131,072
- peak generation tps
- 184
- peak prompt tps
- 2,800
- point count
- 3
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| 1 | Unknown | 113 | 113 | Unknown | Unknown | Unknown | 1 |
| 4 | Unknown | 184 | 46 | Unknown | Unknown | Unknown | 1 |
| 1 | 131,072 | Unknown | Unknown | Unknown | Unknown | 2,800 | 1 |
Provenance & metadata (1)
source
- commit
- 303701cd177e6adecd4999beb67f4797a24439db
- paths
- README.md, Dockerfile, docker-compose.yml, scripts/serve.sh, patches/vllm/models/minimax_m3/nvidia/model.py, patches/vllm/model_executor/layers/quantization/compressed_tensors/compressed_tensors_moe/compressed_tensors_moe_w4a4_mxfp4.py
- repository
- github.com/0xSero/minimax-m3-sm120 ↗