IndexTTS

B
B tier on Text-to-Speech SoftwareScore 7.0· #14of 217
Free plan
Yes
Runs on
api, Linux, self-hosted, Web, Windows
github.com
The IndexTTS homepage

Summary

IndexTTS is ranked #14 of 217 in text-to-speech software on Everything Xiaomi. It runs on API, Linux, Self-hosted, Web, Windows. There is a free plan.

IndexTTS plans and pricing

All plans
IndexTTS Free No hosted plan or price listed; downloadable model and inference code github.com · 9 Oct 2026

Compared on text-to-speech software

Voice cloning
Yesgithub.com
API access
Yesgithub.com
Export formats
WAVgithub.com
Platforms
Web, Windows, Linux, API, self_hostedgithub.com

Facts

Product
IndexTTS is a zero-shot text-to-speech system that clones a voice from a single reference audio clip.github.com · 9 Oct 2026
Languages
IndexTTS-2.5 supports Chinese, English, Japanese, Spanish, and Arabic.github.com · 9 Oct 2026
Voice and emotion controls
The project describes fine-grained emotion control using emotional reference audio, emotion vectors, or text-based emotion input.github.com · 9 Oct 2026
Speaking speed
IndexTTS-2.5 supports duration_factor values from 0.5 to 2.0, with 1.0 as normal speed.github.com · 9 Oct 2026
Pronunciation
IndexTTS-2.5 supports pronunciation control using Chinese Pinyin, English CMU phonemes, and Japanese Kana.github.com · 9 Oct 2026
Deployment
The project documents a local WebUI, a Python API, and production serving through a vLLM recipe.github.com · 9 Oct 2026
Installation platforms
The README gives installation guidance for Windows and Linux and notes that DeepSpeed may be difficult to install on Windows.github.com · 9 Oct 2026
Model downloads
The README provides model download instructions using Hugging Face or ModelScope.github.com · 9 Oct 2026
Hardware
The README recommends NVIDIA CUDA Toolkit 12.8 or newer on Linux or Windows when a CUDA installation error occurs.github.com · 9 Oct 2026
License
The project says it is released under the bilibili Model Use License Agreement and asks users to read its disclaimer before use.github.com · 9 Oct 2026
Commercial use
For commercial usage and cooperation, the project directs users to contact [email protected].github.com · 9 Oct 2026
Official channel and security
The maintainers say the GitHub repository is their only official channel and that they cannot guarantee the security, accuracy, or timeliness of other websites or services.github.com · 9 Oct 2026
Support
The README lists QQ groups, a Discord server, and [email protected] as community contact options.github.com · 9 Oct 2026
Notable pronunciation limit
For IndexTTS-2, Pinyin control works only for supported Chinese Pinyin cases listed in the project's vocabulary file.github.com · 9 Oct 2026
Purpose
IndexTTS is a zero-shot text-to-speech system that clones a voice from a single reference audio clip.github.com · 9 Oct 2026
Emotion control
Speech emotion can be controlled with an emotional reference recording, an emotion vector, or text.github.com · 9 Oct 2026
Speed control
IndexTTS-2.5 supports speaking speed adjustment through duration_factor from 0.5 to 2.0.github.com · 9 Oct 2026
Interfaces
The repository provides a browser-based WebUI and a Python API for inference.github.com · 9 Oct 2026
Requirements
The setup instructions call for Git, uv, and, for Linux or Windows CUDA installations, CUDA Toolkit 12.8 or newer.github.com · 9 Oct 2026
Acceleration
Optional DeepSpeed support may speed up inference on some systems, but the repository says results depend on hardware, drivers, and operating system.github.com · 9 Oct 2026
Official channel
The maintainers identify the GitHub repository as the only official channel maintained by the core team and say other sites or services are not official.github.com · 9 Oct 2026

Best IndexTTS alternatives

See all 20

Where it ranks on Everything Xiaomi

Is IndexTTS yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources