A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
Voice, video, multimodalAUTOMATIC1111
stable-diffusion-webui
The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Who it is for
People who want to generate images locally on their own GPU without writing code.
How to start
- Install the dependencies for your OS (on Debian-based:
sudo apt install wget git python3 python3-venv libgl1 libglib2.0-0). - Run
webui.shon Linux/macOS orwebui-user.baton Windows as a regular user. - Tweak launch options in
webui-user.sh(orwebui-user.bat).
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Stable Diffusion web UI
More in «Voice, video, multimodal»
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.
