The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.
Voice, video, multimodalQwenLM
Qwen3-VL
The Qwen team's family of multimodal models: understand images, documents and video, locate objects in a frame and work with UIs.
Who it is for
People who parse documents, screenshots and video with open models.
How to start
- Try the demo on Hugging Face or Qwen Chat (links at the top of the README).
- Download weights from the Qwen3-VL collection on Hugging Face.
- Open the notebooks in the Cookbooks section: recognition, document parsing, OCR, object grounding.
Steps are taken from the README. Check the current version in the repository before running them.
Stars over the last 30 days
Author's description
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
More in «Voice, video, multimodal»
A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.
OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.
Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.
Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.
A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.
