All repositories

Voice, video, multimodalQwenLM

Qwen3-VL

The Qwen team's family of multimodal models: understand images, documents and video, locate objects in a frame and work with UIs.

Who it is for

People who parse documents, screenshots and video with open models.

How to start

  1. Try the demo on Hugging Face or Qwen Chat (links at the top of the README).
  2. Download weights from the Qwen3-VL collection on Hugging Face.
  3. Open the notebooks in the Cookbooks section: recognition, document parsing, OCR, object grounding.

Steps are taken from the README. Check the current version in the repository before running them.

Stars over the last 30 days

+176Sep 6 — Oct 5
19,86120,037

Author's description

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

The classic browser UI for Stable Diffusion: text-to-image, inpainting, upscaling and thousands of community extensions.

165Kstars+664 in 30 dPython

Comfy-Org

ComfyUI

Growing fastest

A node-based builder for diffusion pipelines: wire up image, video and audio generation visually. Runs locally and exposes API endpoints.

136Kstars+4.7K in 30 dPython

openai

whisper

OpenAI's reference speech-recognition model: transcribes audio in dozens of languages, translates speech and detects the language.

110Kstars+1.6K in 30 dPython

RVC-Boss

GPT-SoVITS

Few-shot voice cloning: about a minute of recorded speech is enough to fine-tune a text-to-speech model. Ships a web UI for data prep and training.

62Kstars+1.1K in 30 dPython

ggml-org

whisper.cpp

Whisper rewritten in C/C++: runs fast on a plain CPU and on Macs with no Python or heavy dependencies. Easy to embed in apps.

54Kstars+759 in 30 dC++

2noise

ChatTTS

A speech-synthesis model tuned for natural conversational dialogue: handles pauses, laughter and intonation. Good for voice assistants and dialogue voice-over.

40Kstars+136 in 30 dPython