The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalComfy-Org
ComfyUI
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
Who it is for
People who need precise control over image and video generation rather than a single button.
How to start
- Download the desktop app from comfy.org/download (Windows and macOS), the easiest way in.
- Or follow the Manual Install section of the README, which supports all OSes and GPU types.
- Open the ready-made templates at comfy.org/workflows and run one as is.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface. The fastest local inference engine in the world.
More in «Voice, video, multimodal»
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.
