Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit
The best open source alternative to llama.cpp is Ollama. If that doesn't suit you, we've compiled a ranked list of other open source llama.cpp alternatives to help you find a suitable replacement. Other interesting open source alternatives to llama.cpp are: LocalAI and Jan.
llama.cpp alternatives are mainly Local Model Runners Tools but may also be AI Personal Assistants or AI Chat Interfaces. Browse these if you want a narrower list of alternatives or looking for a specific functionality of llama.cpp.
Runs open-source language models locally with a simple setup, plus optional cloud access for larger models when local hardware isn't enough.

Ollama lets you run large language models directly on your own hardware, without sending data to third-party APIs. It's built for developers, researchers, and anyone who wants the capabilities of modern AI models without giving up control over their data.
The core idea is local-first. Models run entirely on your machine, which means no usage fees per token, no data leaving your network, and full offline capability for sensitive or mission-critical work. When local hardware isn't enough, an optional cloud tier gives access to larger models running on datacenter-grade hardware.
Key capabilities:
Ollama fits naturally into setups where you want a local AI assistant or a self-hosted backend for agent-based tools. The free tier covers cloud model access at a basic level, with paid plans unlocking higher concurrency and usage limits for heavier workloads.
Run LLMs, speech, image generation, and autonomous agents on your own hardware with an OpenAI-compatible API and 60+ swappable backends.

LocalAI is a self-hosted runtime that lets you run virtually any AI workload on hardware you control. Text generation, vision, speech recognition, text-to-speech, image and video generation, embeddings, reranking, and autonomous agents all run behind a single OpenAI-compatible API. If you're already using OpenAI or Anthropic APIs, switching the endpoint is often all it takes.
The core design is deliberately lean. Backends aren't bundled upfront. They're pulled on demand when a model needs them, each one wrapping a best-in-class engine like llama.cpp, vLLM, SGLang, MLX, or whisper.cpp as an isolated service. You can install, update, or remove individual backends without touching the rest of the stack. Hardware mixing is first-class: NVIDIA, AMD, Intel, Apple Silicon, Vulkan, and Jetson all work, and you can route across them in a single cluster.
For cases where existing engines are too heavy or too closed, the LocalAI team builds its own:
It scales from a CPU-only laptop to a distributed GPU cluster without changing how you interact with it. A single workstation setup can grow into a team server with API keys, roles, quotas, and usage tracking, then further into a multi-worker cluster with model routing and device-spanning inference. Local model runners rarely cover this range in one package.
Agents are built in, not bolted on. You can create agents with MCP tools, memory, RAG, and citations directly from the UI or API. Realtime voice experiences are supported through WebRTC with interruptible STT, LLM output, and TTS pipelines, similar to what LiveKit handles for general media but focused on AI interaction. Privacy controls go beyond keeping data local: PII analysis, redaction middleware, and audit logging are available at the infrastructure level.
The API surface is compatible with OpenAI, Anthropic, Ollama, and ElevenLabs conventions, so existing tooling like LibreChat or Open WebUI connects without custom adapters.
Jan runs open-source AI models on your own hardware or connects to cloud providers like OpenAI, Anthropic, and Google, keeping your conversations off third-party servers.

Jan is a desktop AI chat app built for people who want the capabilities of ChatGPT without handing their conversations to a cloud service. It runs entirely on your own machine using open-source models, or you can connect it to hosted providers when you need more power. Either way, you stay in control.
The model selection is broad. Out of the box, Jan supports:
This flexibility is what separates Jan from simpler local-model runners. You get offline capability when you want privacy, and cloud fallback when a task demands it.
Jan is built in public and has crossed 5.7 million downloads, which puts it among the more widely adopted self-hosted AI interfaces alongside tools like Open WebUI and LibreChat. The codebase is open, community-driven, and free.
The interface is clean and focused on conversation. A memory feature is in development that will carry context and preferences across sessions, so you won't need to re-explain your setup or working style each time you start a new chat.
It's a practical fit for developers, designers, researchers, or anyone who handles sensitive work and doesn't want that data leaving their machine. Local inference does require reasonable hardware, but Jan handles model management inside the app, so you don't need to touch a terminal to get started.
Capture screenshots, generate PDFs, scrape content, extract metadata, and automate browsers with one API.
Get free credits