Recipe

inkling-small-nvfp4-dgxspark-vllm-tp2

inkling-small-nvfp4-dgxspark-vllm-tp2

Saved two-DGX-Spark Inkling Small NVFP4 vLLM profile with MTP1, BF16 KV, SM121 paged-KV overlays, full and piecewise CUDA graphs, and text/image/audio support; blocked from CLI until the host-specific two-node bundle and model revision are independently replayed

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
65b7662d3fcb773afaf751ab29ac6960a0cf011d+sm121-overlays
Graph
full-and-piecewise
Accelerators
2
Tensor parallel
2
Context tokens
262,144
Max concurrency
4
KV cache tokens
420,162
chat
yes
reasoning
yes
tools
yes
vision
yes

Hugging Face model card

Identity

https://huggingface.co/thinkingmachines/Inkling-Small-NVFP4
Repository
thinkingmachines/Inkling-Small-NVFP4
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Compose configuration

vllm

Container · candidate · docker compose

Candidate evidence — not a Run contract

Compose file
compose.yml
Port
8000

Environment

VariableValue
CUTE_DSL_ARCHsm_121a
GPU_MEMORY_UTILIZATION0.9
INKLING_MODEL/models/Inkling-Small-NVFP4
INKLING_VLLM_IMAGElocal/inkling-vllm:20260812-65b7662d-sm121-paged-a80c4d7-audio1
KV_CACHE_DTYPEauto
KV_CACHE_MEMORY_BYTES29450000000
MAX_MODEL_LEN262144
MAX_NUM_BATCHED_TOKENS8192
MAX_NUM_SEQS4
MTP_NUM_TOKENS1
PIPELINE_PARALLEL_SIZE1
SERVED_MODEL_NAMEinkling-small
TENSOR_PARALLEL_SIZE2

Measured speed

ConcurrencyContextPrefillDecodeTTFT msStatusSweep
116,4373,419.7254,806.6historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep
23,482.512.58,187.1historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep
43,5041.813,899.4historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep
15631,683.932.3334.3historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep
22,123.827531.8historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep
42,292.921.1985.6historicalinkling-small-nvfp4-dgxspark-vllm-tp2-sweep

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
2
hardware id
dgx-spark-gb10-128gb
id
inkling-small-nvfp4-dgxspark-vllm-tp2
model instance id
thinkingmachines-inkling-small-nvfp4--nvfp4
recipe source
0xsero
schema version
local-ai-registry/v1
speed sweep ids
inkling-small-nvfp4-dgxspark-vllm-tp2-sweep
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
Yes

engine

graph mode
full-and-piecewise
name
vllm
version
65b7662d3fcb773afaf751ab29ac6960a0cf011d+sm121-overlays

serving

kv cache tokens
420,162
max concurrency
4
max context tokens
262,144
tensor parallel
2
Provenance & metadata (3)

facts

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry