All repositories

Voice, video, multimodalRVC-Boss

GPT-SoVITS

Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.

Who it is for

People who want to narrate text in their own or a chosen voice.

How to start

  1. Create an environment: conda create -n GPTSoVits python=3.10 then conda activate GPTSoVits.
  2. Run the installer: bash install.sh --device <CU126|CU128|ROCM|CPU> --source <HF|HF-Mirror|ModelScope> (on Mac use --device <MPS|CPU>).
  3. Then follow the README: data preparation, fine-tuning and launching the UI.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+1,131Sep 6 — Oct 5
61,22062,351

Author's description

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.

165Kstars+664 in 30 dPython

Comfy-Org

ComfyUI

Growing fastest

A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.

136Kstars+4.7K in 30 dPython

openai

whisper

OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.

110Kstars+1.6K in 30 dPython

ggml-org

whisper.cpp

Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.

54Kstars+759 in 30 dC++

2noise

ChatTTS

A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.

40Kstars+136 in 30 dPython

myshell-ai

OpenVoice

Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.

38Kstars+383 in 30 dPython