A curated collection of the best open source projects tagged "Llm Inference". Each listing includes a website screenshot along with a detailed review of its features.
Run LLMs locally with minimal setup, maximum hardware support
Stars
130,224
Last commit
3 hours ago
License
MIT
C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across a wide range of hardware.
High-throughput, memory-efficient serving engine for LLMs
Stars
93,122
Last commit
4 hours ago
License
Apache-2.0
Inference and serving engine for large language models, built for speed and hardware efficiency with an OpenAI-compatible API and support for a wide range of open models.
Serverless GPU compute with sub-second cold starts
Stars
1,802
Last commit
21 hours ago
License
AGPL-3.0
Run GPU inference, task queues, and sandboxes on serverless infrastructure with sub-second cold starts, autoscaling, and support for your own AWS, GCP, or bare metal.