You cannot select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.
yanyuxiyangzk@126.com 1c8f9338bb vllm文档 1 year ago
..
ChatGPT.py 加入LLM问答,如ChatGPT 2 years ago
Gemini.py 加入LLM问答,如ChatGPT 2 years ago
LLM.py 加入LLM问答,如ChatGPT 2 years ago
Qwen.py 加入LLM问答,如ChatGPT 2 years ago
README.md vllm文档 1 year ago

README.md

1、推理加速 conda create -n vllm python=3.10 conda install pytorch==2.1.0 torchvision==0.16.0 torchaudio==2.1.0 pytorch-cuda=12.1 -c pytorch -c nvidia

python -m vllm.entrypoints.openai.api_server --tensor-parallel-size=1 --trust-remote-code --max-model-len 1024 --model THUDM/chatglm3-6b

python -m vllm.entrypoints.openai.api_server --host 127.0.0.1 --port 8101 --tensor-parallel-size=1 --trust-remote-code --max-model-len 1024 --model THUDM/chatglm3-6b