The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalhexgrad
kokoro
A lightweight 82M-parameter text-to-speech model that sounds far better than its size suggests and runs fast locally.
Who it is for
People who need free voice-over without the cloud or heavy hardware.
How to start
- Install:
pip install kokoro>=0.9.4 soundfile. - Install
espeak-ng(on Linux:apt-get -qq -y install espeak-ng). - Create
KPipeline(lang_code='a')and call it with text and theaf_heartvoice.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Star history is accumulating — the chart appears in a few days.
Author's description
https://hf.co/hexgrad/Kokoro-82M
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
