Recipe

glm-5-2-q3-k-m-rtx-3090-24gb-llama-cpp-tp10

glm-5-2-q3-k-m-rtx-3090-24gb-llama-cpp-tp10

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
llama.cpp
Engine version
9738 (upstream commit e27f30859)
Accelerators
10
Tensor parallel
10
Context tokens
262,144
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF
Repository
pipenetwork/GLM-5.2-REAP50-Q3_K_M-GGUF
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

llama.cpp

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmqqqup2b01a2qo01lwvrysjs

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. llama-server
  2. --ctx-size
  3. 262144
  4. --cache-type-k
  5. q4_0
  6. --cache-type-v
  7. q4_0
  8. --parallel
  9. 1
  10. --flash-attn
  11. on
  12. --split-mode
  13. layer
  14. --tensor-split
  15. 1,1,1,1,1,1,1,1,1,1
  16. --n-gpu-layers
  17. 999
  18. --jinja
  19. --reasoning
  20. off
  21. --no-webui
FlagValue
--ctx-size262144
--cache-type-kq4_0
--cache-type-vq4_0
--parallel1
--flash-attnon
--split-modelayer
--tensor-split1,1,1,1,1,1,1,1,1,1
--n-gpu-layers999
--reasoningoff

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1262,14422.192.2observedglm-5-2-q3-k-m-rtx-3090-24gb-llama-cpp-tp10-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
10
hardware id
rtx-3090-24gb
id
glm-5-2-q3-k-m-rtx-3090-24gb-llama-cpp-tp10
model instance id
pipenetwork-glm-5-2-reap50-q3-k-m-gguf--q3-k-m
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
glm-5-2-q3-k-m-rtx-3090-24gb-llama-cpp-tp10-sweep
status
candidate

capabilities

engine

name
llama.cpp
version
9738 (upstream commit e27f30859)

serving

max concurrency
1
max context tokens
262,144
tensor parallel
10
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
RTX 3090
localmaxxing · notes
GLM-5.2 REAP50 Q3_K_M GGUF on 10x RTX 3090 via llama.cpp CUDA, 262144 context, q4_0 K/V cache, server-side reasoning disabled, single streaming request, 5 measured 1000-token runs after 1 warmup. PCIe 3.0 host, no NVLink. Real-use note: in day-to-day local agent and chat workflows, this GLM-5.2 trial did not perform as well as the restored Qwen 27B coordinator despite the benchmark result.
localmaxxing · observed command
llama-server --ctx-size 262144 --cache-type-k q4_0 --cache-type-v q4_0 --parallel 1 --flash-attn on --split-mode layer --tensor-split 1,1,1,1,1,1,1,1,1,1 --n-gpu-layers 999 --jinja --reasoning off --no-webui
localmaxxing · run id
cmqqqup2b01a2qo01lwvrysjs
localmaxxing · tokenized · arguments
llama-server, --ctx-size, 262144, --cache-type-k, q4_0, --cache-type-v, q4_0, --parallel, 1, --flash-attn, on, --split-mode, layer, --tensor-split, 1,1,1,1,1,1,1,1,1,1, --n-gpu-layers, 999, --jinja, --reasoning, off, --no-webui
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmqqqup2b01a2qo01lwvrysjs