LM/SGLang: Difference between revisions

From Fundamental Ramen
< LM
Jump to navigation Jump to search
No edit summary
 
(9 intermediate revisions by the same user not shown)
Line 1: Line 1:
'''Engine works fine, but most models cannot loaded in xpu mode.'''
= Rebuild docker image from source =
= Rebuild docker image from source =


Line 30: Line 32:


<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
sglang serve --help
sglang serve                       \
 
     --model-path Intel/gemma-4-12B-it-int4-AutoRound \
sglang serve                        \
     --device xpu                   \
     --model-path <MODEL_ID_OR_PATH>  \
     --attention-backend intel_xpu
    --trust-remote-code              \
    --disable-overlap-schedule      \
     --device xpu                     \
    --host 0.0.0.0                  \
    --tp 1                          \  # using multi GPUs
     --attention-backend intel_xpu   \  # using intel optimized XPU attention backend
    --page-size                      \  # intel_xpu attention backend supports [32, 64, 128]
</syntaxhighlight>
</syntaxhighlight>



Latest revision as of 09:31, 17 September 2026

Engine works fine, but most models cannot loaded in xpu mode.

Rebuild docker image from source

cd sglang
git pull
cd docker
docker build -t sglang-xpu:latest -f xpu.Dockerfile .

Run

docker run \
    -it \
    --privileged \
    --ipc=host \
    --network=host \
    --user root \
    --group-add $(getent group video | cut -d: -f3) \
    --device /dev/dri \
    -v /dev/dri/by-path:/dev/dri/by-path \
    -v /dev/shm:/dev/shm \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 30000:30000 \
    -e "HF_TOKEN=$HF_TOKEN" \
    sglang-xpu:latest /bin/bash

In container ...

sglang serve                        \
    --model-path Intel/gemma-4-12B-it-int4-AutoRound \
    --device xpu                    \
    --attention-backend intel_xpu

Structure

TODO