Recipe

gemma-4-12b-q4-k-m-dgx-spark-gb10-128gb-lmstudio-tp1

gemma-4-12b-q4-k-m-dgx-spark-gb10-128gb-lmstudio-tp1

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
lmstudio
Engine version
2.20.1
Accelerators
1
Tensor parallel
1
Context tokens
66,861
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/google/gemma-4-12B
Repository
google/gemma-4-12B
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

lmstudio

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmqdwa20m00n2mp01f41powhb

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. LM
  2. Studio
  3. runtime
  4. llama.cpp-linux-arm64-nvidia-cuda13
  5. v2.20.1;
  6. model
  7. google/gemma-4-12b
  8. Q4_K_M;
  9. ctx
  10. 66861;
  11. temp
  12. 0;
  13. max_tokens
  14. 512

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
166,86125.8167observedgemma-4-12b-q4-k-m-dgx-spark-gb10-128gb-lmstudio-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
gemma-4-12b-q4-k-m-dgx-spark-gb10-128gb-lmstudio-tp1
model instance id
google-gemma-4-12b--q4-k-m
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
gemma-4-12b-q4-k-m-dgx-spark-gb10-128gb-lmstudio-tp1-sweep
status
candidate

capabilities

engine

name
lmstudio
version
2.20.1

serving

max concurrency
1
max context tokens
66,861
tensor parallel
1
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:26:09Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:26:09Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:26:09Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:26:09Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:26:09Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:26:09Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
GB10 Grace Blackwell
localmaxxing · notes
Benchmarked via LM Studio REST /api/v0/chat/completions. Median of 2 measured 512-token generations after 1 warm-up; temperature 0; fixed prompt (43 prompt tokens). Runtime: llama.cpp-linux-arm64-nvidia-cuda13 v2.20.1, CUDA 13.0, driver 580.142. NVIDIA DGX Spark (GB10 Grace Blackwell), 20-core Grace ARM (Cortex-X925/A725), 121 GiB unified memory. Note: gemma-4 is a reasoning model; ~509 of 512 output tokens were reasoning tokens.
localmaxxing · observed command
LM Studio runtime llama.cpp-linux-arm64-nvidia-cuda13 v2.20.1; model google/gemma-4-12b Q4_K_M; ctx 66861; temp 0; max_tokens 512
localmaxxing · run id
cmqdwa20m00n2mp01f41powhb
localmaxxing · tokenized · arguments
LM, Studio, runtime, llama.cpp-linux-arm64-nvidia-cuda13, v2.20.1;, model, google/gemma-4-12b, Q4_K_M;, ctx, 66861;, temp, 0;, max_tokens, 512
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:26:09Z

sources

captured atkindurl
2026-08-30T09:26:09Znormalized-recipewww.localmaxxing.com/en/runs/cmqdwa20m00n2mp01f41powhb