The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalblack-forest-labs
flux
The official inference code for FLUX.1 models: text-to-image generation and editing, including Kontext mode where you edit an image with words.
Who it is for
People who want to run FLUX on their own GPU or inside their own code.
How to start
- Clone:
cd $HOME && git clone https://github.com/black-forest-labs/fluxandcd $HOME/flux. - Create an environment with
python3.10 -m venv .venvand install:pip install -e ".[all]". - For image editing run
python -m flux kontext --track_usage --loop(requiresBFL_API_KEY).
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Official inference repo for FLUX.1 models
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
