High-throughput, memory-efficient serving engine for LLMs
Stars
93,122
Last commit
8 hours ago
License
Apache-2.0
Inference and serving engine for large language models, built for speed and hardware efficiency with an OpenAI-compatible API and support for a wide range of open models.