Speed sweep

nvidia-nemotron-3-5-lightning-30b-a3b-bf16-q4-0-intel-arc-pro-b70-32gb-llama-cpp-tp1-sweep

nvidia-nemotron-3-5-lightning-30b-a3b-bf16-q4-0-intel-arc-pro-b70-32gb-llama-cpp-tp1-sweep

Record

Recipe
nvidia-nemotron-3-5-lightning-30b-a3b-bf16-q4-0-intel-arc-pro-b70-32gb-llama-cpp-tp1
Measured
2026-08-12T06:53:16.931Z
Points
1

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatus
116,38476687.4observed

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

id
nvidia-nemotron-3-5-lightning-30b-a3b-bf16-q4-0-intel-arc-pro-b70-32gb-llama-cpp-tp1-sweep
measured at
2026-08-12T06:53:16.931Z
recipe id
nvidia-nemotron-3-5-lightning-30b-a3b-bf16-q4-0-intel-arc-pro-b70-32gb-llama-cpp-tp1
schema version
local-ai-registry/v1

metrics

concurrency
1
inference engine version
b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch
latest point at
2026-08-12T06:53:16.931Z
max context tokens
16,384
peak generation tps
87.4
peak prompt tps
766
point count
1

rows

concurrencycontext tokensdecode tok sdecode tok s per streamoutput tokenspeak vram gbprefill tok ssamples
116,38487.487.40Unknown7661
Provenance & metadata (1)