Skip to main content

Inference Engines and Versions

The available minor versions, container images, and driver requirements of the inference engines built into the CSGHub platform are summarized below, for reference when creating frameworks and deploying instances.

Built-in Inference Engines and Versions​

NVIDIA GPU​

Choose the entry matching the CUDA version supported by the cluster driver.

EngineVersionImageCUDAModel formatUse case
vllmv0.28.0opencsghq/vllm:v0.28.013.0safetensorsText generation / embedding / reranking / speech
vllmv0.9.2opencsghq/vllm:v0.9.2-cu11811.8safetensorsText generation
nvidia-vllm25.11-py3opencsghq/nvidia-vllm:25.11-py313.0safetensorsText generation / embedding / reranking
sglangv0.5.14opencsghq/sglang:v0.5.14-cu13013.0safetensorsText generation
nvidia-sglang25.12-py3opencsghq/nvidia-sglang:25.12-py313.0safetensorsText generation
sglang-qwen3-guard-streamv0.5.3rc0opencsghq/sglang:0.5.3rc0-qwen3-guard-stream12.8.1safetensorsContent moderation (Qwen3-Guard streaming)
tgi3.2.1opencsghq/tgi:3.2.112.4safetensorsText generation
tei1.7opencsghq/tei:1.712.2safetensorsEmbedding / reranking
llama.cppb9787opencsghq/llama.cpp:b978712.4ggufText generation
ktransformers0.2.1.post1opencsghq/ktransformers:0.2.1.post112.4ggufText generation
hf-inference-toolkit0.5.5opencsghq/hf-inference-toolkit:0.5.512.8safetensorsText-to-image
diffusers0.39.0opencsghq/diffusers:0.39.012.8safetensorsText-to-image / image editing
diffusers0.38.0opencsghq/diffusers:0.38.012.8safetensorsText-to-image / image editing (legacy)
funasr1.3.9opencsghq/funasr:1.3.912.8pytorchSpeech recognition
paddleocr3.7.0opencsghq/paddleocr:3.7.012.8paddle_staticOCR
paddleocr-vl3.7.0opencsghq/paddleocr-vl:3.7.012.8safetensorsOCR (vision-language)
audiofly1.0opencsghq/audiofly:1.012.1pytorchText-to-audio
audio-fish1.5.1opencsghq/fish-speech:server-cuda12.6pytorchSpeech synthesis (fish-speech)
lightx2v0.0.1opencsghq/lightx2v:26011201-cu12812.8safetensorsVideo generation
longcat-video1.5opencsghq/longcat-video:1.5-cu12412.4safetensorsAudio-driven avatar video

CPU​

No driver version requirement.

EngineVersionImageModel formatUse case
vllmv0.24.0opencsghq/vllm-cpu:v0.24.0safetensorsText generation
tei1.7opencsghq/tei:cpu-1.7safetensorsEmbedding / reranking
llama.cppb9787opencsghq/llama.cpp:b9787-cpuggufText generation
funasr1.3.9opencsghq/funasr-cpu:1.3.9pytorchSpeech recognition
paddleocr3.7.0opencsghq/paddleocr-cpu:3.7.0paddle_staticOCR
audio-fish1.5.1opencsghq/fish-speech:server-cpupytorchSpeech synthesis (fish-speech)

AMD ROCm GPU​

The driver version is the ROCm version the image is built for.

EngineVersionImageROCmModel formatUse case
amd-vllm0.28.0opencsghq/amd-vllm:rocm7.2.1_vllm_0.28.07.2.1safetensorsText generation / embedding / reranking
amd-tgi3.3.6opencsghq/tgi-rocm:rocm6.3.1_tgi_3.3.66.3.1safetensorsText generation
amd-llama.cppb9787opencsghq/llama.cpp-rocm:rocm7.2.2-b97877.2.2ggufText generation
amd-funasr1.3.9opencsghq/funasr-rocm:1.3.97.2.2pytorchSpeech recognition
amd-diffusers0.39.0opencsghq/diffusers-rocm:0.39.07.2.2safetensorsText-to-image / image editing
amd-diffusers0.38.0opencsghq/diffusers-rocm:0.38.07.2.2safetensorsText-to-image / image editing (legacy)
amd-hf-inference-toolkit0.5.5opencsghq/hf-inference-toolkit-rocm:0.5.57.2.2safetensorsText-to-image
amd-audiofly1.0opencsghq/audiofly-rocm:1.07.2.2pytorchText-to-audio
amd-longcat-video1.5opencsghq/longcat-video-rocm:1.5-rocm7.2.27.2.2safetensorsAudio-driven avatar video

Huawei Ascend NPU​

Driver versions are bundled in the images.

EngineVersionImageModel formatUse case
ascend-vllmv0.25.1rcopencsghq/ascend-vllm:v0.25.1rcsafetensorsText generation
mindie1.0.RC2opencsghq/mindie:2.0-csg-1.0.RC2safetensorsText generation

MetaX GPGPU​

Driver versions are bundled in the images.

EngineVersionImageModel formatUse case
vllmv0.25.0opencsghq/metax-vllm:0.25.0safetensorsText generation
sglangv0.5.12opencsghq/metax-sglang:0.5.12safetensorsText generation

Hygon DCU​

The driver version is the DTK version the image is built for.

EngineVersionImageDTKModel formatUse case
vllmv0.8.5opencsghq/vllm:v0.8.5-dtk25.0425.04safetensorsText generation

Kunlun GCU​

See the table below for driver versions.

EngineVersionImageDriverModel formatUse case
tgi-gcu3.2.1opencsghq/tgi-gcu:tgi_25.02.22.125.02safetensorsText generation