All repositories

Open models & inferenceDao-AILab

flash-attention

A fast, memory-efficient implementation of attention that large-model training and inference on GPUs lean on. The result is exact, with no approximation.

Who it is for

For people speeding up transformer training and inference on NVIDIA GPUs.

How to start

  1. Install packaging, psutil and ninja, which the build needs.
  2. Install: pip install flash-attn --no-build-isolation.
  3. Call flash_attn_func(q, k, v, causal=True) instead of standard attention; on Hopper and Blackwell there is pip install flash-attn-4.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+269Sep 6 — Oct 5
24,81925,088

Author's description

Fast and memory-efficient exact attention

ollama

ollama

The easiest way to run an open model on your own machine: one command downloads and starts it, plus a local REST API and libraries for Python and JavaScript.

182Kstars+2.6K in 30 dGo

huggingface

transformers

The core library for working with models: one codebase for text, vision, audio and multimodal tasks, for both inference and training. Most new open models ship through it.

167Kstars+2.4K in 30 dPython

open-webui

open-webui

A self-hosted ChatGPT-style interface that connects to Ollama and any OpenAI-compatible API. Installs with a single Docker command and runs on your own server.

154Kstars+3.3K in 30 dPython

ggml-org

llama.cpp

A C/C++ engine that runs language models on ordinary hardware: laptops, phones, GPU-less servers. Much of local AI, Ollama included, is built on it.

130Kstars+3.6K in 30 dC++

deepseek-ai

DeepSeek-V3

Repository of the large open DeepSeek-V3 mixture-of-experts model: description, benchmark results and run instructions. It showed an open model can stand next to closed ones.

105Kstars+327 in 30 dPython

pytorch

pytorch

The foundation under most modern neural networks: GPU-accelerated tensors and automatic differentiation, wrapped in ordinary Python. Transformers, Llama and nearly every open model are written on it.

104Kstars+1.1K in 30 dPython