On this page
sglang0.5.11-ascend-cann8.5.0
Prerequisites
- Architecture: aarch64
- Chip models: Ascend 910C
- Host driver: 25.5.0
- Container toolkit: Ascend-docker-runtime
Image contents
Python
3.11
Application package
sglang==0.5.11
sglang-fl==0.2.0
Component versions
| Component | Version |
|---|---|
| FlagGems | 5.4.0 (6ed2f390) |
| FlagTree | 0.6.0+ascend3.2 |
| FlagCX | 0.13.0 |
| torch | 2.8.0 (torch_npu 2.8.0.post2) |
| CANN | 8.5.0 |
Environment
ASCEND_RT_VISIBLE_DEVICES— which devices the server uses, e.g.0,1,2,3FLAGCX_PATH=/sgl-workspace/FlagCXHCCL_BUFFSIZE=1000HCCL_OP_EXPANSION_MODE=AIVSTREAMS_PER_DEVICE=32SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1SGLANG_NPU_USE_MULTI_STREAM=1
The CANN and ATB environments must be sourced before launching:
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
Launch
Published: harbor.baai.ac.cn/flagos-app/sglang0.5.11-ascend-cann8.5.0:2.2.0-0.2.0
The image name is long — assign it to a variable first:
IMG=harbor.baai.ac.cn/flagos-app/sglang0.5.11-ascend-cann8.5.0:2.2.0-0.2.0
This image ships no default entrypoint command, and there is no sglang-serve wrapper in it — start the server with python3 -m sglang.launch_server as shown below.
Start an interactive shell with the container toolkit:
docker run --rm -it \
--runtime ascend \
--privileged \
--network host \
--shm-size 32g \
-e ASCEND_VISIBLE_DEVICES=0,1,2,3 \
-v /path/to/models:/models \
$IMG bash
Then, inside the container, source the environments and launch the server:
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
export FLAGCX_PATH=/sgl-workspace/FlagCX
export HCCL_BUFFSIZE=1000
export HCCL_OP_EXPANSION_MODE=AIV
export STREAMS_PER_DEVICE=32
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
export SGLANG_NPU_USE_MULTI_STREAM=1
python3 -m sglang.launch_server \
--model-path /models/Qwen3.6-35B-A3B --tokenizer-path /models/Qwen3.6-35B-A3B \
--host 0.0.0.0 --port 30000 --tp 4 \
--attention-backend ascend --device npu --dtype bfloat16 \
--context-length 32768 --mem-fraction-static 0.8 \
--cuda-graph-max-bs 60 --disable-radix-cache --trust-remote-code
Notes:
--tpmust match the number of devices you expose. Devices are selected withASCEND_VISIBLE_DEVICESatdocker runtime andASCEND_RT_VISIBLE_DEVICESinside the container.- The server needs a large
/dev/shm; a container created without--shm-sizegets 64 MB and the server hangs during startup. - The blacklist of FlagGems operators that are known not to compile or run on this stack is shipped inside the image (
sglang_fl/dispatch/config/ascend.yaml); it is applied automatically, nothing has to be exported. - For a thinking model (e.g. the Qwen3 family), add
--reasoning-parser qwen3to separate the chain of thought intoreasoning_content.
Last updated 24 Sep 2026, 18:29 +0800.