The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalRVC-Boss
GPT-SoVITS
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Who it is for
People who want to narrate text in their own or a chosen voice.
How to start
- Create an environment:
conda create -n GPTSoVits python=3.10thenconda activate GPTSoVits. - Run the installer:
bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope>(on Mac use--device <MPS|CPU>). - Then follow the README: data preparation, fine-tuning and launching the UI.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.
