The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalggml-org
whisper.cpp
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
Who it is for
People who need offline speech transcription on a laptop or built into a product.
How to start
- Clone the repo:
git clone https://github.com/ggml-org/whisper.cpp.gitandcd whisper.cpp. - Download a model:
sh ./models/download-ggml-model.sh base.en. - Build with
cmake -B buildandcmake --build build -j --config Release, then run./build/bin/whisper-cli -f samples/jfk.wav.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Port of OpenAI's Whisper model in C/C++
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
Instant voice cloning: takes a short sample, transfers the timbre to any text and lets you control style, emotion and language. MIT-licensed and free for commercial use.
