Recipe

gemma-4-26b-a4b-it-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

gemma-4-26b-a4b-it-nvfp4-dgx-spark-gb10-128gb-vllm-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
see Agentic Arcade source metadata
Accelerators
1
Tensor parallel
1
Context tokens
2,048
Max concurrency
64
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/nvidia/Gemma-4-26B-A4B-NVFP4
Repository
nvidia/Gemma-4-26B-A4B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmr82nlmd005kqr01o87h22b9

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. Agentic
  2. Arcade
  3. recorded
  4. throughput
  5. row
  6. from
  7. /home/frosty40/agentic-arcade/data/results_throughput_gemma.csv
  8. at
  9. agents=64

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
642,0486,449.4807.35,520observedgemma-4-26b-a4b-it-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
gemma-4-26b-a4b-it-nvfp4-dgx-spark-gb10-128gb-vllm-tp1
model instance id
nvidia-gemma-4-26b-a4b-nvfp4--nvfp4
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
gemma-4-26b-a4b-it-nvfp4-dgx-spark-gb10-128gb-vllm-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
--model, nvidia/Gemma-4-26B-A4B-NVFP4, --tensor-parallel-size, 1, --host, 0.0.0.0, --port, 8000, --max-model-len, 2048
container port
8,000
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm
version
see Agentic Arcade source metadata

serving

max concurrency
64
max context tokens
2,048
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
GB10 Grace Blackwell
localmaxxing · notes
Full DGX/GB10 Agentic Arcade throughput queue. This queue includes the complete missing curve rows rather than silently culling single-stream, warmup-candidate, startup-outlier, or lowest-curve cells. qualityFlags=none. Imported from Agentic Arcade /home/frosty40/agentic-arcade/data/results_throughput_gemma.csv row agents=64. Hardware/model fields inferred from build_site.py where not explicit; run API dry-run before public submit.
localmaxxing · observed command
Agentic Arcade recorded throughput row from /home/frosty40/agentic-arcade/data/results_throughput_gemma.csv at agents=64
localmaxxing · run id
cmr82nlmd005kqr01o87h22b9
localmaxxing · tokenized · arguments
Agentic, Arcade, recorded, throughput, row, from, /home/frosty40/agentic-arcade/data/results_throughput_gemma.csv, at, agents=64
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:26:09Z

sources

captured atkindurl
2026-08-30T09:26:09Znormalized-recipewww.localmaxxing.com/en/runs/cmr82nlmd005kqr01o87h22b9