LM/SGLang: Difference between revisions

From Fundamental Ramen
< LM
Jump to navigation Jump to search
No edit summary
 
(One intermediate revision by the same user not shown)
Line 1: Line 1:
'''Engine works fine, but most models cannot loaded in xpu mode.'''
= Rebuild docker image from source =
= Rebuild docker image from source =


Line 31: Line 33:
<syntaxhighlight lang="bash">
<syntaxhighlight lang="bash">
sglang serve                        \
sglang serve                        \
     --model-path nicosuter/Qwen3.8-27B-AWQ \
     --model-path Intel/gemma-4-12B-it-int4-AutoRound \
     --device xpu                    \
     --device xpu                    \
     --attention-backend intel_xpu
     --attention-backend intel_xpu

Latest revision as of 09:31, 17 September 2026

Engine works fine, but most models cannot loaded in xpu mode.

Rebuild docker image from source

cd sglang
git pull
cd docker
docker build -t sglang-xpu:latest -f xpu.Dockerfile .

Run

docker run \
    -it \
    --privileged \
    --ipc=host \
    --network=host \
    --user root \
    --group-add $(getent group video | cut -d: -f3) \
    --device /dev/dri \
    -v /dev/dri/by-path:/dev/dri/by-path \
    -v /dev/shm:/dev/shm \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    -p 30000:30000 \
    -e "HF_TOKEN=$HF_TOKEN" \
    sglang-xpu:latest /bin/bash

In container ...

sglang serve                        \
    --model-path Intel/gemma-4-12B-it-int4-AutoRound \
    --device xpu                    \
    --attention-backend intel_xpu

Structure

TODO