All repositories

Voice, video, multimodalblack-forest-labs

flux

The official inference code for FLUX.1 models: text-to-image generation and editing, including Kontext mode where you edit an image with words.

Who it is for

People who want to run FLUX on their own GPU or inside their own code.

How to start

  1. Clone: cd $HOME && git clone https://github.com/black-forest-labs/flux and cd $HOME/flux.
  2. Create an environment with python3.10 -m venv .venv and install: pip install -e ".[all]".
  3. For image editing run python -m flux kontext --track_usage --loop (requires BFL_API_KEY).

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+113Sep 6 — Oct 5
25,89526,008

Author's description

Official inference repo for FLUX.1 models

The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.

165Kstars+664 in 30 dPython

Comfy-Org

ComfyUI

Growing fastest

A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.

136Kstars+4.7K in 30 dPython

openai

whisper

OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.

110Kstars+1.6K in 30 dPython

RVC-Boss

GPT-SoVITS

Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.

62Kstars+1.1K in 30 dPython

ggml-org

whisper.cpp

Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.

54Kstars+759 in 30 dC++

2noise

ChatTTS

A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.

40Kstars+136 in 30 dPython