Qwen3-TTS

C
C tier on Text-to-Speech SoftwareScore 5.9 · #151 of 215
Android app
Not listed
Free plan
Yes
Runs on
api, Linux, self-hosted, Web
github.com
The Qwen3-TTS homepage

Summary

Qwen3-TTS is ranked #151 of 215 in text-to-speech software on Everything Xiaomi. It runs on API, Linux, Self-hosted, Web. There is a free plan.

Qwen3-TTS plans and pricing

All plans
Qwen3-TTS open-source models Free The project is released under the Apache-2.0 license; no price is stated. 0.6B and 1.7B models · 10 languages github.com · 7 Oct 2026

Compared on text-to-speech software

Commercial use
Yesgithub.com
Voice cloning
Yesgithub.com
API access
Yesgithub.com
Languages
10 languagesgithub.com
Export formats
WAVgithub.com
Platforms
Web, API, self_hostedgithub.com

Facts

Product
Qwen3-TTS is an open-source series of text-to-speech models developed by the Qwen team at Alibaba Cloud.github.com · 7 Oct 2026
Generation modes
The models support streaming and non-streaming speech generation, voice design, voice cloning, and natural-language voice control.github.com · 7 Oct 2026
Latency
The README says the streaming architecture can produce the first audio packet after one character and reports end-to-end synthesis latency as low as 97 ms.github.com · 7 Oct 2026
Model sizes
Released models include 0.6B and 1.7B parameter variants for custom voice or base tasks, plus a 1.7B voice design model.github.com · 7 Oct 2026
Local installation
The project provides a Python package install and says it recommends a fresh Python 3.12 environment.github.com · 7 Oct 2026
Local web interface
A local Gradio web UI demo can be launched with the qwen-tts-demo command.github.com · 7 Oct 2026
Model hosting
The README links model downloads on Hugging Face and ModelScope and documents loading weights locally or by model ID.github.com · 7 Oct 2026
API integration
The project links Alibaba Cloud DashScope real-time APIs for custom voice, voice cloning, and voice design.github.com · 7 Oct 2026
Inference integration
The README states that vLLM-Omni supports offline Qwen3-TTS inference and that online serving is planned for later support.github.com · 7 Oct 2026
License
The GitHub repository lists the Apache-2.0 license.github.com · 7 Oct 2026
API requirements
The linked Alibaba Cloud real-time synthesis guide says API use requires configuring an API key and installing the latest DashScope SDK when using that SDK.help.aliyun.com · 7 Oct 2026
What it does
Qwen3-TTS is an open-source series of text-to-speech models for speech generation, voice design, and voice cloning.github.com · 7 Oct 2026
Voice options
The released models support custom voices, voice design from natural-language descriptions, and voice cloning from reference audio.github.com · 7 Oct 2026
Streaming
The README says the models support streaming and non-streaming generation and reports end-to-end synthesis latency as low as 97 ms.github.com · 7 Oct 2026
Voice cloning input
Voice cloning uses reference audio and its transcript; the README says using only the speaker embedding without a transcript may reduce cloning quality.github.com · 7 Oct 2026
Download and install
Models can be downloaded through Hugging Face or ModelScope, and the project provides a `qwen-tts` Python package installable from PyPI.github.com · 7 Oct 2026
Hardware and setup
The README demonstrates model loading on CUDA and recommends FlashAttention 2 to reduce GPU memory use, subject to compatible hardware.github.com · 7 Oct 2026
Web demo
The project includes a local web UI demo launched with `qwen-tts-demo`.github.com · 7 Oct 2026
Serving limitation
The README says vLLM-Omni supports offline inference currently, with online serving planned for later.github.com · 7 Oct 2026
API pricing
Alibaba Cloud lists international Qwen3-TTS API rates of $0.115 per 10,000 input characters for Voice Design and Voice Cloning, with output free; the page also lists a 110,000-character free quota for each.alibabacloud.com · 7 Oct 2026

Best Qwen3-TTS alternatives

See all 20

Where it ranks on Everything Xiaomi

Is Qwen3-TTS yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources