Sign inSign up

scitrera/dgx-spark-sglang

By scitrera

Updated about 1 month ago

sglang container image optimized for NVIDIA DGX Spark

Image
5

10K+

scitrera/dgx-spark-sglang repository overview

CUDA Containers for NVIDIA DGX Spark

https://github.com/scitrera/cuda-containers

This repository contains Dockerfiles and build recipes for CUDA-based containers optimized for NVIDIA DGX Spark systems, with a focus on vLLM, sglang, PyTorch, and multi-node inference workloads.

The primary goal of this project is to provide stable, well-versioned, prebuilt images that work out-of-the-box on DGX Spark (Blackwell-ready), while still being suitable as base images for custom builds.


Why This Repo Exists

The official NVIDIA images tend to run too far behind the latest releases. Other community images prioritize bleeding edge over versioning and stability.

The goal of this repo is to provide a stable, well-versioned, prebuilt images that work out-of-the-box on DGX Spark (Blackwell-ready).

The main architectural difference from other builds (e.g. eugr's repo (link below) -- which is pretty much the community standard) is:

  • NCCL and PyTorch are built first, in a dedicated base image
  • vLLM and related tooling are layered on top
  • Versioning follows vLLM releases as the primary axis

For sglang, the officially provided container is not continuously updated. I assume that might change in the near future as sglang gets better SM121 support -- but in the meantime, Scitrera will, on a best effort basis, maintain sglang images similar to our vLLM images.


Available Images

SGLang Images

SGLang images are also optimized for DGX Spark and provide an alternative high-performance inference runtime.

Latest Releases
SGLang 0.5.8
  • scitrera/dgx-spark-sglang:0.5.8-t4

    • SGLang 0.5.8 (with build fixes post-release)
    • PyTorch 2.10.0 (with torchvision + torchaudio)
    • CUDA 13.1.1
    • Transformers 4.57.6
    • Triton 3.6.0
    • NCCL 2.29.3-1
    • FlashInfer 0.6.3
  • scitrera/dgx-spark-sglang:0.5.8-t5

    • Same as above, but with Transformers 5.1.0

PyTorch Development Base Image

If you want to build your own inference stack:

  • scitrera/dgx-spark-pytorch-dev:2.10.0-v2-cu131

    • PyTorch 2.10.0
    • CUDA 13.1.1
    • NCCL 2.29.3-1
    • Built on nvidia/cuda:13.1.1-devel-ubuntu24.04
    • Includes standard build tooling
  • scitrera/dgx-spark-pytorch-dev:2.10.0-cu131

    • PyTorch 2.10.0
    • CUDA 13.1.0
    • NCCL 2.29.2-1
    • Built on nvidia/cuda:13.1.0-devel-ubuntu24.04
    • Includes standard build tooling

This is the recommended base image if you want to:

  • Build vLLM/sglang/other tools yourself
  • Add custom kernels or extensions
  • Experiment with alternative runtimes

Tag Semantics

Tags follow this pattern for vLLM and SGLang containers:


<version>-t<transformers-major>

Examples:

  • 0.13.0-t4 → vLLM 0.13.0 + Transformers 4.x
  • 0.5.8-t5 → SGLang 0.5.8 + Transformers 5.x

Example Usage (SGLang)

docker run \
  --privileged \
  --gpus all \
  -it --rm \
  --network host --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  scitrera/dgx-spark-sglang:0.5.8-t4 \
  sglang serve \
    --model-path Qwen/Qwen2.5-7B-Instruct \
    --mem-fraction-static 0.4

Inspecting Component Versions

Major component versions are embedded as Docker labels.

docker inspect scitrera/dgx-spark-vllm:0.14.0rc2-t4 \
  --format '{{json .Config.Labels}}' | jq

Example output:

{
  "dev.scitrera.cuda_version": "13.1.0",
  "dev.scitrera.flashinfer_version": "0.6.1",
  "dev.scitrera.nccl_version": "2.28.9-1",
  "dev.scitrera.torch_version": "2.10.0-rc6",
  "dev.scitrera.transformers_version": "4.57.5",
  "dev.scitrera.triton_version": "3.5.1",
  "dev.scitrera.vllm_version": "0.14.0rc2"
}

Notes & Caveats

  • NCCL is upgraded relative to upstream PyTorch builds
  • PyTorch, Triton, and vLLM/sglang are rebuilt accordingly
  • Image sizes could still be optimized further
  • Version combinations are chosen to be as new as possible but limited by stability (not guaranteed to have the latest features if they might break things)

Roadmap (Loose)

  • Better size optimization
  • More documentation/support for DGX Spark newcomers

This project is not affiliated with NVIDIA. This project is sponsored and maintained by scitrera.ai.

Tag summary

Content type

Image

Digest

sha256:cc1cec4d0

Size

16.7 GB

Last updated

about 1 month ago

docker pull scitrera/dgx-spark-sglang:0.5.17