The easiest way to run an open model on your own machine: one command downloads and starts it, plus a local REST API and libraries for Python and JavaScript.
Open models & inferencemudler
LocalAI
A local replacement for cloud APIs: an OpenAI-compatible server that runs text, voice and image models on any hardware, no GPU required.
Who it is for
For people who want a full model API at home without the cloud.
How to start
- Run it in Docker:
docker run -ti --name local-ai -p 8080:8080 localai/localai:latest. - Load a model from the gallery:
local-ai run llama-3.2-1b-instruct:q4_k_m. - Chat from the terminal:
local-ai chat --model llama-3.2-1b-instruct:q4_k_m.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
More in «Open models & inference»
The core library for working with models: one codebase for text, vision, audio and multimodal tasks, for both inference and training. Most new open models ship through it.
A self-hosted ChatGPT-style interface that connects to Ollama and any OpenAI-compatible API. Installs with a single Docker command and runs on your own server.
A C/C++ engine that runs language models on ordinary hardware: laptops, phones, GPU-less servers. Much of local AI, Ollama included, is built on it.
Repository of the large open DeepSeek-V3 mixture-of-experts model: description, benchmark results and run instructions. It showed an open model can stand next to closed ones.
The foundation under most modern neural networks: GPU-accelerated tensors and automatic differentiation, wrapped in ordinary Python. Transformers, Llama and nearly every open model are written on it.
