Recipe

step-3-7-flash-ud-iq4-nl-gguf-dgx-spark-gb10-128gb-llama-cpp-tp1

step-3-7-flash-ud-iq4-nl-gguf-dgx-spark-gb10-128gb-llama-cpp-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
see Agentic Arcade source metadata
Accelerators
1
Tensor parallel
1
Context tokens
2,048
Max concurrency
3
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/stepfun-ai/Step-3.7-Flash-GGUF
Repository
stepfun-ai/Step-3.7-Flash-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmr84wguo007oqr01kgd5lcmk

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. Agentic
  2. Arcade
  3. recorded
  4. throughput
  5. row
  6. from
  7. /home/frosty40/agentic-arcade/data/results_throughput_step37.csv
  8. at
  9. agents=3

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
32,04894.211.220,140observedstep-3-7-flash-ud-iq4-nl-gguf-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
step-3-7-flash-ud-iq4-nl-gguf-dgx-spark-gb10-128gb-llama-cpp-tp1
model instance id
stepfun-ai-step-3-7-flash-gguf--ud-iq4-nl-gguf
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
step-3-7-flash-ud-iq4-nl-gguf-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
-hf, stepfun-ai/Step-3.7-Flash-GGUF, --n-gpu-layers, 999, --host, 0.0.0.0, --port, 8080, -c, 2048
container port
8,080
environment · LLAMA CACHE
/root/.cache/huggingface
host port
8,080
image
ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107ff48e4720145911c86930f2ce
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
gemma-4-12b-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
synthesized · template
llama-cpp-server-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
llama.cpp
version
see Agentic Arcade source metadata

serving

max concurrency
3
max context tokens
2,048
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
draft launch.entrypoint · provenance · captured at
2026-09-01T01:41:26Z
draft launch.entrypoint · reason linux-amd64-container-config-entrypoint
draft launch.entrypoint · state known
engine.graph mode · provenance · captured at
2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-31T23:03:15Z
serving.max concurrency · reason server-capacity-derived-from-source-evidence
serving.max concurrency · state known

metadata

localmaxxing · backend
cuda
localmaxxing · batch size
3
localmaxxing · hardware label
GB10 Grace Blackwell
localmaxxing · notes
Full DGX/GB10 Agentic Arcade throughput queue. This queue includes the complete missing curve rows rather than silently culling single-stream, warmup-candidate, startup-outlier, or lowest-curve cells. qualityFlags=metadata_corrected_from_source_doc. Imported from Agentic Arcade /home/frosty40/agentic-arcade/data/results_throughput_step37.csv row agents=3. Hardware/model fields inferred from build_site.py where not explicit; run API dry-run before public submit.
localmaxxing · observed command
Agentic Arcade recorded throughput row from /home/frosty40/agentic-arcade/data/results_throughput_step37.csv at agents=3
localmaxxing · run id
cmr84wguo007oqr01kgd5lcmk
localmaxxing · tokenized · arguments
Agentic, Arcade, recorded, throughput, row, from, /home/frosty40/agentic-arcade/data/results_throughput_step37.csv, at, agents=3
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:26:09Z

sources

captured atkindurl
2026-08-30T09:26:09Znormalized-recipewww.localmaxxing.com/en/runs/cmr84wguo007oqr01kgd5lcmk