Recipe
qwen38-27b-nvfp4-dgxspark-sglang-tp1
qwen38-27b-nvfp4-dgxspark-sglang-tp1MiaAI-Lab Qwen3.8-27B NVFP4 single-Spark DSpark profile, normalized to preserve CUDA graphs and awaiting index-protocol revalidation
Record
- Status
- candidate
- Source
- mialabs
- Engine
- sglang
- Engine version
- qwen38-27b image
- Graph
- piecewise
- Accelerators
- 1
- Tensor parallel
- 1
- Max concurrency
- 16
- chat
- yes
- reasoning
- yes
- tools
- yes
- vision
- yes
Hugging Face model card
Identity
https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4- Repository
- RadixArk/Qwen3.8-27B-NVFP4
- Status
- known
- Link type
- Exact Hub repository
Public Hugging Face repository confirmed by the Hub API.
Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.
Docker configuration
sglang
Container · candidate · docker
Candidate evidence — not a Run contract
- Image
lmsysorg/sglang:qwen38-27b@sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1- Digest
sha256:febfb971c7352570fc445c466ebd6ffc9d896024958e544a60f2137fd85856b1- Port
- 8888
Observed source tokens
Mechanical split of the source command. Unverified against the engine CLI. The registry does not offer Run for this recipe.
python3-msglang.launch_server--model-pathRadixArk/Qwen3.8-27B-NVFP4--revision319f741cce68d7914884900c138a1fbb70a42f30--served-model-nameqwen3.8-27b--trust-remote-code--mem-fraction-static0.90--attention-backendflashinfer--chunked-prefill-size8192--kv-cache-dtypefp8_e4m3--mamba-ssm-dtypebfloat16--mamba-full-memory-ratio4.21--mamba-radix-cache-strategyextra_buffer_lazy--max-mamba-cache-size64--max-running-requests16--context-length262144--speculative-algorithmDSPARK--speculative-draft-model-pathRadixArk/Qwen3.8-27B-DSpark--speculative-dspark-block-size7--speculative-draft-model-quantizationunquant--speculative-num-draft-tokens8--enable-torch-compile--torch-compile-max-bs4--cuda-graph-max-bs-decode4--num-continuous-decode-steps2--reasoning-parserqwen3--tool-call-parserqwen3_coder--enable-metrics--enable-cache-report--host0.0.0.0--port8888
| Flag | Value |
|---|---|
-m | sglang.launch_server |
--model-path | RadixArk/Qwen3.8-27B-NVFP4 |
--revision | 319f741cce68d7914884900c138a1fbb70a42f30 |
--served-model-name | qwen3.8-27b |
--mem-fraction-static | 0.90 |
--attention-backend | flashinfer |
--chunked-prefill-size | 8192 |
--kv-cache-dtype | fp8_e4m3 |
--mamba-ssm-dtype | bfloat16 |
--mamba-full-memory-ratio | 4.21 |
--mamba-radix-cache-strategy | extra_buffer_lazy |
--max-mamba-cache-size | 64 |
--max-running-requests | 16 |
--context-length | 262144 |
--speculative-algorithm | DSPARK |
--speculative-draft-model-path | RadixArk/Qwen3.8-27B-DSpark |
--speculative-dspark-block-size | 7 |
--speculative-draft-model-quantization | unquant |
--speculative-num-draft-tokens | 8 |
--torch-compile-max-bs | 4 |
--cuda-graph-max-bs-decode | 4 |
--num-continuous-decode-steps | 2 |
--reasoning-parser | qwen3 |
--tool-call-parser | qwen3_coder |
--host | 0.0.0.0 |
--port | 8888 |
Environment
| Variable | Value |
|---|---|
SGLANG_OPT_MAMBA_SKIP_DECODE_LOCK | 0 |
Mounts
| Source | Target |
|---|---|
~/.cache/huggingface | /root/.cache/huggingface |
Measured speed
| Concurrency | Context | Prefill | Decode | TTFT ms | Status | Sweep |
|---|---|---|---|---|---|---|
| 1 | — | — | 51.5 | — | historical | qwen38-27b-nvfp4-dgxspark-sglang-tp1-sweep |
Remaining fields
Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.
- hardware count
- 1
- hardware id
- dgx-spark-gb10-128gb
- id
- qwen38-27b-nvfp4-dgxspark-sglang-tp1
- model instance id
- radixark-qwen3-8-27b-nvfp4--nvfp4
- recipe source
- mialabs
- schema version
- local-ai-registry/v1
- speed sweep ids
- qwen38-27b-nvfp4-dgxspark-sglang-tp1-sweep
- status
- candidate
capabilities
- chat
- Yes
- reasoning
- Yes
- tools
- Yes
- vision
- Yes
engine
- graph mode
- piecewise
- name
- sglang
- version
- qwen38-27b image
serving
- max concurrency
- 16
- tensor parallel
- 1
Provenance & metadata (3)
facts
- serving.kv cache tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
- serving.max context tokens · provenance · captured at
- 2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
metadata
provenance
- captured at
- 2026-08-27T06:04:13.773Z
sources
| captured at | kind | url |
|---|---|---|
| 2026-08-27T06:04:13.773Z | normalized-recipe | github.com/0xSero/local-ai-registry ↗ |