Recipe
laguna-s-2-1-q4-k-m-dgx-spark-gb10-128gb-llama-cpp-tp1
laguna-s-2-1-q4-k-m-dgx-spark-gb10-128gb-llama-cpp-tp1Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.
Record
- Status
- candidate
- Source
- localmaxxing
- Engine
- llama.cpp
- Engine version
- 1 (04b2b72)
- Accelerators
- 1
- Tensor parallel
- 1
- Context tokens
- 9,280
- Max concurrency
- 1
- chat
- unknown
- reasoning
- unknown
- tools
- unknown
- vision
- unknown
Hugging Face model card
Identity
https://huggingface.co/poolside/Laguna-S-2.1-GGUF- Repository
- poolside/Laguna-S-2.1-GGUF
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Observed configuration
llama.cpp
Evidence only · candidate · reference
Candidate evidence — not a Run contract
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
llama-server--modellaguna-s-2.1-Q4_K_M.gguf--aliaspoolside/Laguna-S-2.1-GGUF--ctx-size9280--parallel1--batch-size2048--ubatch-size1024--n-gpu-layersauto--fiton--fit-target24576--flash-attnon--cache-type-kq4_0--cache-type-vq4_0--cache-ram0--cache-reuse0--context-checkpoints0--host127.0.0.1--port8000--metrics
| Flag | Value |
|---|---|
--model | laguna-s-2.1-Q4_K_M.gguf |
--alias | poolside/Laguna-S-2.1-GGUF |
--ctx-size | 9280 |
--parallel | 1 |
--batch-size | 2048 |
--ubatch-size | 1024 |
--n-gpu-layers | auto |
--fit | on |
--fit-target | 24576 |
--flash-attn | on |
--cache-type-k | q4_0 |
--cache-type-v | q4_0 |
--cache-ram | 0 |
--cache-reuse | 0 |
--context-checkpoints | 0 |
--host | 127.0.0.1 |
--port | 8000 |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | 9,280 | 785.3 | 19.3 | 10,431.6 | observed | laguna-s-2-1-q4-k-m-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- dgx-spark-gb10-128gb
- id
- laguna-s-2-1-q4-k-m-dgx-spark-gb10-128gb-llama-cpp-tp1
- model instance id
- poolside-laguna-s-2-1-gguf--q4-k-m
- recipe source
- localmaxxing
- schema version
- local-ai-registry/v1
- speed sweep ids
- laguna-s-2-1-q4-k-m-dgx-spark-gb10-128gb-llama-cpp-tp1-sweep
- status
- candidate
capabilities
draft launch
- accelerator backend
- nvidia
- arguments
- -hf, poolside/Laguna-S-2.1-GGUF:Q4_K_M, --n-gpu-layers, 999, --host, 0.0.0.0, --port, 8080, -c, 9280
- container port
- 8,080
- environment · LLAMA CACHE
- /root/.cache/huggingface
- host port
- 8,080
- image
- ghcr.io/ggml-org/llama.cpp:server-cuda12-b10481@sha256:b2497f8834f5ecb4e38530f6bf2734b8e0be107f0f0857e259672d1cb85b71c2
- ipc
- host
- kind
- docker
- shm size
- 16g
- synthesized · generated at
- 2026-08-31T22:12:17Z
- synthesized · image provenance
- gemma-4-12b-q4-k-m-rtx-3060-12gb-llama-cpp-tp1
- synthesized · template
- llama-cpp-server-v1
mounts
| read only | target |
|---|---|
| No | /root/.cache/huggingface |
engine
- name
- llama.cpp
- version
- 1 (04b2b72)
serving
- max concurrency
- 1
- max context tokens
- 9,280
- tensor parallel
- 1
Provenance & metadata (3)
facts
- capabilities.chat · provenance · captured at
- 2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
- capabilities.reasoning · provenance · captured at
- 2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
- capabilities.tools · provenance · captured at
- 2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
- capabilities.vision · provenance · captured at
- 2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
- engine.graph mode · provenance · captured at
- 2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
- serving.kv cache tokens · provenance · captured at
- 2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
metadata
- localmaxxing · backend
- cuda
- localmaxxing · hardware label
- GB10 Grace Blackwell
- localmaxxing · notes
- Corrected ordinary-decoding GGUF result on one NVIDIA DGX Spark GB10. Seven distinct semantic prompts were measured once each with no warmup or prefix reuse. Six prompts contained 8,192 endpoint tokens and one contained 8,190; every request generated 1,024 tokens at concurrency one. Decode samples were 19.2, 19.3, 19.3, 19.3, 19.4, 19.4, and 19.4 tok/s; median 19.3. Fresh TTFT samples were 10,897.33, 10,419.64, 10,418.99, 10,462.31, 10,454.06, 10,376.58, and 10,431.60 ms; median 10,431.60 ms. Effective fresh-prefill samples were 751.7, 786.2, 786.3, 782.8, 783.6, 789.5, and 785.3 prompt tok/s; median 785.3. Poolside Laguna llama.cpp fork 04b2b72 served Q4_K_M revision e08e1fe855bb2d43f96ad78e24495283f3426c67, artifact SHA-256 7da520c5f44bc3c79d4eeebfd1151ba7114c5d7568e72a995638417093c5753f. Stable configuration: 9,280-token declared capacity, one slot, Q4_0 K/V cache, batch 2,048, microbatch 1,024, FlashAttention, all 49 layers offloaded, cache RAM disabled, prompt caching disabled, context checkpoints disabled, and no speculation. Minimum available RAM was 41.59 GiB and minimum free swap was 7.54 GiB; no guard event occurred. The old 18,155.28 ms result was real but used batch and microbatch 256, forcing 32 prompt-processing chunks, and left server prompt-cache/checkpoint behavior enabled. Prefill here is prompt tokens divided by fresh, client-observed TTFT, not kernel-only throughput. Rejected candidates and guard events are archived.
- localmaxxing · observed command
- llama-server --model laguna-s-2.1-Q4_K_M.gguf --alias poolside/Laguna-S-2.1-GGUF --ctx-size 9280 --parallel 1 --batch-size 2048 --ubatch-size 1024 --n-gpu-layers auto --fit on --fit-target 24576 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-ram 0 --cache-reuse 0 --context-checkpoints 0 --host 127.0.0.1 --port 8000 --metrics
- localmaxxing · run id
- cmryoj31g0435o401cbxy7hen
- localmaxxing · tokenized · arguments
- llama-server, --model, laguna-s-2.1-Q4_K_M.gguf, --alias, poolside/Laguna-S-2.1-GGUF, --ctx-size, 9280, --parallel, 1, --batch-size, 2048, --ubatch-size, 1024, --n-gpu-layers, auto, --fit, on, --fit-target, 24576, --flash-attn, on, --cache-type-k, q4_0, --cache-type-v, q4_0, --cache-ram, 0, --cache-reuse, 0, --context-checkpoints, 0, --host, 127.0.0.1, --port, 8000, --metrics
- localmaxxing · tokenized · fidelity
- faithful
provenance
- captured at
- 2026-08-30T09:26:09Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-30T09:26:09Z | normalized-recipe | www.localmaxxing.com/en/runs/cmryoj31g0435o401cbxy7hen ↗ |