~ / skills / devops / serving-llms-vllm

serving-llms-vllm

Serves LLMs with high throughput using vLLM's PagedAttention and continuous batching. Use when deploying production LLM APIs, optimizing inference latency/throughput, or serving…

Industry
devops
License
Unverified
Source repo
Orchestra-Research/AI-Research-SKILLs · ★ 10,718
Source file
12-inference-serving/vllm/SKILL.md
View full SKILL.md on GitHub →

不会安装?看中文图文教程 →