Recipe

glm-5-3-flash-awq-4bit-rtx-3090-24gb-vllm-tp8

glm-5-3-flash-awq-4bit-rtx-3090-24gb-vllm-tp8

Observed LocalMaxxing leaderboard run. Evidence for compatibility, not an executable launch contract.

Record

Status
candidate
Source
localmaxxing
Engine
vllm
Engine version
0.28.1rc1.dev125+g85872952a.glm53sm86
Accelerators
8
Tensor parallel
8
Context tokens
204,800
Max concurrency
1
chat
unknown
reasoning
unknown
tools
unknown
vision
unknown

Hugging Face model card

Identity

https://huggingface.co/wtdcode/GLM-5.3-Flash-AWQ-W4A16
Repository
wtdcode/GLM-5.3-Flash-AWQ-W4A16
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Observed configuration

vllm

Evidence only · candidate · reference

Candidate evidence — not a Run contract

Source
https://www.localmaxxing.com/en/runs/cmtebgn9q006clm01eray6kbm

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. /home/thread/venvs/vllm-glm53-sm86-85872952/bin/python
  2. /home/thread/venvs/vllm-glm53-sm86-85872952/bin/vllm
  3. serve
  4. /mnt/nvme/models/wtdcode/GLM-5.3-Flash-AWQ-W4A16-target-only
  5. --served-model-name
  6. glm-5.3-flash-awq-sm86-test
  7. --host
  8. 127.0.0.1
  9. --port
  10. 18083
  11. --tensor-parallel-size
  12. 8
  13. --pipeline-parallel-size
  14. 1
  15. --distributed-executor-backend
  16. mp
  17. --enable-expert-parallel
  18. --enable-ep-weight-filter
  19. --all2all-backend
  20. allgather_reducescatter
  21. --disable-custom-all-reduce
  22. --language-model-only
  23. --load-format
  24. safetensors
  25. --safetensors-load-strategy
  26. lazy
  27. --max-parallel-loading-workers
  28. 1
  29. --attention-backend
  30. TRITON_MLA_SPARSE
  31. --dtype
  32. bfloat16
  33. --kv-cache-dtype
  34. fp8
  35. --max-model-len
  36. 204800
  37. --max-num-seqs
  38. 1
  39. --max-num-batched-tokens
  40. 256
  41. --enable-chunked-prefill
  42. --block-size
  43. 128
  44. --kv-cache-memory-bytes
  45. 1298373120
  46. --kernel-config
  47. {moe_backend:marlin,enable_flashinfer_autotune:false}
  48. --compilation-config
  49. {mode:VLLM_COMPILE,cudagraph_mode:PIECEWISE,cudagraph_capture_sizes:[1],max_cudagraph_capture_size:1}
  50. --enable-auto-tool-choice
  51. --tool-call-parser
  52. glm47
  53. --reasoning-parser
  54. glm47
FlagValue
--served-model-nameglm-5.3-flash-awq-sm86-test
--host127.0.0.1
--port18083
--tensor-parallel-size8
--pipeline-parallel-size1
--distributed-executor-backendmp
--all2all-backendallgather_reducescatter
--load-formatsafetensors
--safetensors-load-strategylazy
--max-parallel-loading-workers1
--attention-backendTRITON_MLA_SPARSE
--dtypebfloat16
--kv-cache-dtypefp8
--max-model-len204800
--max-num-seqs1
--max-num-batched-tokens256
--block-size128
--kv-cache-memory-bytes1298373120
--kernel-config{moe_backend:marlin,enable_flashinfer_autotune:false}
--compilation-config{mode:VLLM_COMPILE,cudagraph_mode:PIECEWISE,cudagraph_capture_sizes:[1],max_cudagraph_capture_size:1}
--tool-call-parserglm47
--reasoning-parserglm47

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
1204,8001,02075.2357.8observedglm-5-3-flash-awq-4bit-rtx-3090-24gb-vllm-tp8-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
8
hardware id
rtx-3090-24gb
id
glm-5-3-flash-awq-4bit-rtx-3090-24gb-vllm-tp8
model instance id
wtdcode-glm-5-3-flash-awq-w4a16--awq-4bit
recipe source
localmaxxing
schema version
local-ai-registry/v1
speed sweep ids
glm-5-3-flash-awq-4bit-rtx-3090-24gb-vllm-tp8-sweep
status
candidate

capabilities

draft launch

accelerator backend
nvidia
arguments
--model, wtdcode/GLM-5.3-Flash-AWQ-W4A16, --tensor-parallel-size, 8, --host, 0.0.0.0, --port, 8000, --max-model-len, 204800
container port
8,000
host port
8,000
image
vllm/vllm-openai@sha256:0a51ea5b4ae2dc5d81890e5173f54203d2a3ae0cfffe51b8fd2afd4391bfd967
ipc
host
kind
docker
shm size
16g
synthesized · generated at
2026-08-31T22:12:17Z
synthesized · image provenance
deepseek-fp8-rtx-pro-6000-blackwell-96gb-vllm-tp1
synthesized · template
vllm-openai-v1

mounts

read onlytarget
No/root/.cache/huggingface

engine

name
vllm
version
0.28.1rc1.dev125+g85872952a.glm53sm86

serving

max concurrency
1
max context tokens
204,800
tensor parallel
8
Provenance & metadata (3)

facts

capabilities.chat · provenance · captured at
2026-08-30T09:10:02Z
capabilities.chat · reason capability-not-verified
capabilities.chat · state unknown
capabilities.reasoning · provenance · captured at
2026-08-30T09:10:02Z
capabilities.reasoning · reason capability-not-verified
capabilities.reasoning · state unknown
capabilities.tools · provenance · captured at
2026-08-30T09:10:02Z
capabilities.tools · reason capability-not-verified
capabilities.tools · state unknown
capabilities.vision · provenance · captured at
2026-08-30T09:10:02Z
capabilities.vision · reason capability-not-verified
capabilities.vision · state unknown
engine.graph mode · provenance · captured at
2026-08-30T09:10:02Z
engine.graph mode · reason runtime-detail-not-published
engine.graph mode · state unknown
serving.kv cache tokens · provenance · captured at
2026-08-30T09:10:02Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown

metadata

localmaxxing · backend
cuda
localmaxxing · hardware label
RTX 3090
localmaxxing · notes
AWQ W4A16 routed-expert checkpoint at HF revision ac8c3e52fd4f; local target-only view omits optional NextN draft shards, so speculative decoding is off. Custom SM86 vLLM fork, TP8+EP8, Triton sparse MLA, Marlin MoE, FP8 KV, piecewise CUDA graphs. Two 64-token warmups and six sequential 512-token measurements; temperature 0, seed 260829, ignore_eos=true. Prefix caching is enabled by the engine; each request used a unique cache_salt and every Prometheus cached-token/hit delta was zero. tokSOut is client steady-state; tokSTotal=(prompt+output)/client E2E per current API wording. Temperature-zero hashes: 6 unique/6; all six runs exhausted the 512-token budget in reasoning_content before visible content. PID/config stable=True; restarts unchanged=True; runner incidents=0; benchmark-window service/kernel suspicious lines=0/0. Per-GPU power is median of each run's mean active draw; peak VRAM is summed across all GPUs.
localmaxxing · observed command
/home/thread/venvs/vllm-glm53-sm86-85872952/bin/python /home/thread/venvs/vllm-glm53-sm86-85872952/bin/vllm serve /mnt/nvme/models/wtdcode/GLM-5.3-Flash-AWQ-W4A16-target-only --served-model-name glm-5.3-flash-awq-sm86-test --host 127.0.0.1 --port 18083 --tensor-parallel-size 8 --pipeline-parallel-size 1 --distributed-executor-backend mp --enable-expert-parallel --enable-ep-weight-filter --all2all-backend allgather_reducescatter --disable-custom-all-reduce --language-model-only --load-format safetensors --safetensors-load-strategy lazy --max-parallel-loading-workers 1 --attention-backend TRITON_MLA_SPARSE --dtype bfloat16 --kv-cache-dtype fp8 --max-model-len 204800 --max-num-seqs 1 --max-num-batched-tokens 256 --enable-chunked-prefill --block-size 128 --kv-cache-memory-bytes 1298373120 --kernel-config {"moe_backend":"marlin","enable_flashinfer_autotune":false} --compilation-config {"mode":"VLLM_COMPILE","cudagraph_mode":"PIECEWISE","cudagraph_capture_sizes":[1],"max_cudagraph_capture_size":1} --enable-auto-tool-choice --tool-call-parser glm47 --reasoning-parser glm47
localmaxxing · run id
cmtebgn9q006clm01eray6kbm
localmaxxing · tokenized · arguments
/home/thread/venvs/vllm-glm53-sm86-85872952/bin/python, /home/thread/venvs/vllm-glm53-sm86-85872952/bin/vllm, serve, /mnt/nvme/models/wtdcode/GLM-5.3-Flash-AWQ-W4A16-target-only, --served-model-name, glm-5.3-flash-awq-sm86-test, --host, 127.0.0.1, --port, 18083, --tensor-parallel-size, 8, --pipeline-parallel-size, 1, --distributed-executor-backend, mp, --enable-expert-parallel, --enable-ep-weight-filter, --all2all-backend, allgather_reducescatter, --disable-custom-all-reduce, --language-model-only, --load-format, safetensors, --safetensors-load-strategy, lazy, --max-parallel-loading-workers, 1, --attention-backend, TRITON_MLA_SPARSE, --dtype, bfloat16, --kv-cache-dtype, fp8, --max-model-len, 204800, --max-num-seqs, 1, --max-num-batched-tokens, 256, --enable-chunked-prefill, --block-size, 128, --kv-cache-memory-bytes, 1298373120, --kernel-config, {moe_backend:marlin,enable_flashinfer_autotune:false}, --compilation-config, {mode:VLLM_COMPILE,cudagraph_mode:PIECEWISE,cudagraph_capture_sizes:[1],max_cudagraph_capture_size:1}, --enable-auto-tool-choice, --tool-call-parser, glm47, --reasoning-parser, glm47
localmaxxing · tokenized · fidelity
faithful

provenance

captured at
2026-08-30T09:10:02Z

sources

captured atkindurl
2026-08-30T09:10:02Znormalized-recipewww.localmaxxing.com/en/runs/cmtebgn9q006clm01eray6kbm