The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalfacebookresearch
sam2
Meta's Segment Anything 2: picks out any object in an image or video from a click and tracks it across frames. The repo has code, checkpoints and notebooks.
Who it is for
People doing annotation, video editing and computer vision.
How to start
- Install:
git clone https://github.com/facebookresearch/sam2.git && cd sam2, thenpip install -e .. - For notebooks add:
pip install -e ".[notebooks]". - Download checkpoints:
cd checkpoints && ./download_ckpts.sh && cd ...
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
