Recipe

hy3-fp8-rtxpro6000-vllm-tp4

hy3-fp8-rtxpro6000-vllm-tp4

Controller-backed candidate captured read-only from Pop!_OS. Saved Pop controller config; candidate pending inference-index acceptance.

Record

Status
candidate
Source
0xsero
Engine
vllm
Engine version
controller-managed; captured 2026-08-23
Graph
unknown
Accelerators
4
Tensor parallel
4
chat
yes
reasoning
yes
tools
yes
vision
no

Hugging Face model card

Identity

https://huggingface.co/tencent/Hy3-FP8
Repository
tencent/Hy3-FP8
Status
known
Link type
Exact Hub repository

Public Hugging Face repository confirmed by the Hub API.

Candidate: useful compatibility or speed evidence. The registry does not offer Run until promotion requirements are met.

Controller configuration

vllm

No container · candidate · controller

Candidate evidence — not a Run contract

Environment

VariableValue
CUDA_DEVICE_ORDERPCI_BUS_ID
CUDA_VISIBLE_DEVICESGPU-fa982c97-64af-db6a-2ffb-08380e1f9375,GPU-c6ac75f2-cadf-6ff3-4cab-76c6033c1006,GPU-3cece4bc-432e-705e-7324-3f441d9cb4cc,GPU-7b5db8b3-3a49-c0a7-b8e4-80dc1bd3e853
FLASHINFER_DISABLE_VERSION_CHECK1
GLOO_SOCKET_IFNAMElo
HF_HOME${MODEL_ROOT}/hf
LD_LIBRARY_PATH${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cu13/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cublas/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cuda_cupti/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cuda_nvrtc/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cuda_runtime/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cudnn/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cufft/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cufile/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/curand/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cusolver/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cusparse/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/cusparselt/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/nccl/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/nvjitlink/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/nvshmem/lib:${HOME}/.venvs/vllm-0.22.0/lib/python3.10/site-packages/nvidia/nvtx/lib:
NCCL_IB_DISABLE1
NCCL_P2P_DISABLE0
NCCL_P2P_LEVELPIX
NCCL_SOCKET_IFNAMElo
OMP_NUM_THREADS8
PYTORCH_CUDA_ALLOC_CONFexpandable_segments:True
SAFETENSORS_FAST_GPU1
TORCH_NCCL_ASYNC_ERROR_HANDLING1
TORCH_NCCL_BLOCKING_WAIT1
VLLM_USE_V11

Remaining fields

Identity, launch, related records, and measured speed are shown above. This is the rest of the normalized record.

hardware count
4
hardware id
rtx-pro-6000-blackwell-96gb
id
hy3-fp8-rtxpro6000-vllm-tp4
model instance id
tencent-hy3-fp8--fp8
recipe source
0xsero
schema version
local-ai-registry/v1
status
candidate

capabilities

chat
Yes
reasoning
Yes
tools
Yes
vision
No

engine

graph mode
unknown
name
vllm
version
controller-managed; captured 2026-08-23

serving

tensor parallel
4
Provenance & metadata (3)

facts

serving.kv cache tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.kv cache tokens · reason kv-cache-capacity-not-published
serving.kv cache tokens · state unknown
serving.max concurrency · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max concurrency · reason not-observed
serving.max concurrency · state unknown
serving.max context tokens · provenance · captured at
2026-08-27T06:04:13.773Z
serving.max context tokens · reason not-observed
serving.max context tokens · state unknown
speed sweep ids · provenance · captured at
2026-08-27T06:04:13.773Z
speed sweep ids · reason not-observed
speed sweep ids · state unknown

metadata

provenance

captured at
2026-08-27T06:04:13.773Z

sources

captured atkindurl
2026-08-27T06:04:13.773Znormalized-recipegithub.com/0xSero/local-ai-registry