Any one can run this model with SGlang framework？

#13

by muziyongshixin - opened Feb 20

Feb 20

•

I try to run this model with SGlang, but it is extremely slow. Does anyone have a good setting to run this model with SGlang?

v2ray

Cognitive Computations org Feb 20

Try run this with vLLM, it is much faster.

Eric108

Feb 21

you can try this command python3 -m sglang.launch_server --host 0.0.0.0 --port 30000 --model-path models/DeepSeek-R1-AWQ --tp 8 --enable-p2p-check --trust-remote-code --dtype float16 --mem-fraction-static 0.95 --served-model-name deepseek-r1-awq --disable-cuda-graph with sglang==0.4.2
but the result is not as expected..I get empty content for some queries and the think is not complete

oliver0102

19 days ago

I managed to run this model with sglang
python3 -m sglang.launch_server --model-path /home/service/var/models/deepseek-r1-huggingface/DeepSeek-R1-AWQ/ --trust-remote-code --tp 8 --mem-fraction-static 0.8 --dtype float16 --host 0.0.0.0 --port 9000 --disable-radix --disable-custom-all-reduce --log-requests --cuda-graph-max-bs 16 --max-total-tokens 65536. It runs on my 8*H20 with 30t/s. However i have same issue as @Eric10 . Tried some parameter but finally give up. I am now trying to use q4 gguf model.

isofun

4 days ago

My decoding running much more slowly with sglang 0.4.4 than fp8 and there are always some strange output.

v2ray

Cognitive Computations org 3 days ago

What's the reason that you must use SGLang instead of vLLM, since vLLM got all the features now, MLA, MTP, a mush faster fused MoE Marlin kernel, etc..

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

Your need to confirm your account before you can post a new comment.

· Sign up or log in to comment