Speed sweep
nvidia-nemotron-3-5-gptq-int4-g64-sym-local-dflash-bf16-local-intel-arc-pro-b70-32gb-vllm-tp1-sweep
nvidia-nemotron-3-5-gptq-int4-g64-sym-local-dflash-bf16-local-intel-arc-pro-b70-32gb-vllm-tp1-sweepRecord
- Measured
- 2026-08-13
- Points
- 1
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status |
|---|---|---|---|---|---|
| — | 16,384 | 7,160 | 186.6 | — | historical |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- id
- nvidia-nemotron-3-5-gptq-int4-g64-sym-local-dflash-bf16-local-intel-arc-pro-b70-32gb-vllm-tp1-sweep
- measured at
- 2026-08-13
- recipe id
- nvidia-nemotron-3-5-gptq-int4-g64-sym-local-dflash-bf16-local-intel-arc-pro-b70-32gb-vllm-tp1
- schema version
- local-ai-registry/v1
metrics
- inference engine version
- v0.26.1rc1.dev668+g3ee2df303 (XPU)
- latest point at
- 2026-08-13
- max context tokens
- 16,384
- peak generation tps
- 186.6
- peak prompt tps
- 7,160
- point count
- 1
rows
| concurrency | context tokens | decode tok s | decode tok s per stream | output tokens | peak vram gb | prefill tok s | samples |
|---|---|---|---|---|---|---|---|
| Unknown | 16,384 | 186.6 | Unknown | 0 | Unknown | 7,160 | 1 |
Provenance & metadata (1)
source
- repository
- www.localmaxxing.com ↗