Recipe

qwen38-27b-nvfp4-dgxspark-sglang-tp1

qwen38-27b-nvfp4-dgxspark-sglang-tp1

MiaAI-Lab Qwen3.8-27B NVFP4 single-Spark DSpark profile, normalized to preserve CUDA graphs and awaiting index-protocol revalidation

Record

Status
candidate
Source
mialabs
Engine
sglang
Engine version
qwen38-27b image
Graph
piecewise
Accelerators
1
Tensor parallel
1
Max concurrency
16
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4
Repository
RadixArk/Qwen3.8-27B-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Docker configuration

sglang

Container · candidate · docker

Candidate evidence — not a Run contract

Image
lmsysorg/sglang:qwen38-27b@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
Digest
sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1
Port
8888

Observed source tokens

Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.

  1. python3
  2. -m
  3. sglang.launch_server
  4. --model-path
  5. RadixArk/Qwen3.8-27B-NVFP4
  6. --revision
  7. 319f741cce68d7914884900c138a1fbb70a42f30
  8. --served-model-name
  9. qwen3.8-27b
  10. --trust-remote-code
  11. --mem-fraction-static
  12. 0.90
  13. --attention-backend
  14. flashinfer
  15. --chunked-prefill-size
  16. 8192
  17. --kv-cache-dtype
  18. fp8_e4m3
  19. --mamba-ssm-dtype
  20. bfloat16
  21. --mamba-full-memory-ratio
  22. 4.21
  23. --mamba-radix-cache-strategy
  24. extra_buffer_lazy
  25. --max-mamba-cache-size
  26. 64
  27. --max-running-requests
  28. 16
  29. --context-length
  30. 262144
  31. --speculative-algorithm
  32. DSPARK
  33. --speculative-draft-model-path
  34. RadixArk/Qwen3.8-27B-DSpark
  35. --speculative-dspark-block-size
  36. 7
  37. --speculative-draft-model-quantization
  38. unquant
  39. --speculative-num-draft-tokens
  40. 8
  41. --enable-torch-compile
  42. --torch-compile-max-bs
  43. 4
  44. --cuda-graph-max-bs-decode
  45. 4
  46. --num-continuous-decode-steps
  47. 2
  48. --reasoning-parser
  49. qwen3
  50. --tool-call-parser
  51. qwen3_coder
  52. --enable-metrics
  53. --enable-cache-report
  54. --host
  55. 0.0.0.0
  56. --port
  57. 8888
FlagValue
-msglang.launch_server
--model-pathRadixArk/Qwen3.8-27B-NVFP4
--revision319f741cce68d7914884900c138a1fbb70a42f30
--served-model-nameqwen3.8-27b
--mem-fraction-static0.90
--attention-backendflashinfer
--chunked-prefill-size8192
--kv-cache-dtypefp8_e4m3
--mamba-ssm-dtypebfloat16
--mamba-full-memory-ratio4.21
--mamba-radix-cache-strategyextra_buffer_lazy
--max-mamba-cache-size64
--max-running-requests16
--context-length262144
--speculative-algorithmDSPARK
--speculative-draft-model-pathRadixArk/Qwen3.8-27B-DSpark
--speculative-dspark-block-size7
--speculative-draft-model-quantizationunquant
--speculative-num-draft-tokens8
--torch-compile-max-bs4
--cuda-graph-max-bs-decode4
--num-continuous-decode-steps2
--reasoning-parserqwen3
--tool-call-parserqwen3_coder
--host0.0.0.0
--port8888

Environment

VariableValue
SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK0

Mounts

SourceTarget
~/.cache/huggingface/root/.cache/huggingface

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
151.5historicalqwen38-27b-nvfp4-dgxspark-sglang-tp1-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
1
hardware id
dgx-spark-gb10-128gb
id
qwen38-27b-nvfp4-dgxspark-sglang-tp1
model instance id
radixark-qwen3-8-27b-nvfp4--nvfp4
recipe source
mialabs
schema version
local-ai-registry/v1
speed sweep ids
qwen38-27b-nvfp4-dgxspark-sglang-tp1-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
piecewise
name
sglang
version
qwen38-27b image

serving

max concurrency
16
tensor parallel
1
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry