Run LLMs locally with minimal setup, maximum hardware support
Stars
124,853
Last commit
11 hours ago
License
MIT
C/C++ inference engine for large language models, supporting quantization, multi-GPU, Apple Silicon, and an OpenAI-compatible server across a wide range of hardware.
The purpose-built time series platform for high-velocity ingestion and real-time queries at scale, without sacrificing performance or cost. Download InfluxDB for free.