IndexTTS-2
About IndexTTS-2
IndexTTS2 is a breakthrough auto-regressive zero-shot Text-to-Speech (TTS) model designed to address the challenge of precise duration control in large-scale TTS systems, which is a significant limitation in applications like video dubbing. It introduces a novel, general method for speech duration control, supporting two modes: one that explicitly specifies the number of generated tokens for precise duration, and another that generates speech freely in an auto-regressive manner. Furthermore, IndexTTS2 achieves disentanglement between emotional expression and speaker identity, enabling independent control over timbre and emotion via separate prompts. To enhance speech clarity in highly emotional expressions, the model incorporates GPT latent representations and utilizes a novel three-stage training paradigm. To lower the barrier for emotional control, it also features a soft instruction mechanism based on text descriptions, developed by fine-tuning Qwen3, to effectively guide the generation of speech with the desired emotional tone. Experimental results show that IndexTTS2 outperforms state-of-the-art zero-shot TTS models in word error rate, speaker similarity, and emotional fidelity across multiple datasets
Available Serverless
Run queries immediately, pay only for usage
$
7.15
Per 1M UTF-8 bytes
Metadata
Specification
State
Available
Architecture
Calibrated
Yes
Mixture of Experts
No
Total Parameters
1
Activated Parameters
Reasoning
No
Precision
FP8
Context length
0K
Max Tokens
Supported Functionality
Serverless
Supported
Serverless LoRA
Not supported
Fine-tuning
Not supported
Embeddings
Not supported
Rerankers
Not supported
Support image input
Not supported
JSON Mode
Not supported
Structured Outputs
Not supported
Tools
Not supported
Fim Completion
Not supported
Chat Prefix Completion
Not supported
Compare with Other Models
See how this model stacks up against others.
IndexTeam
text-to-speech
IndexTTS-2
Release on: Sep 10, 2025
Total Context:
0K
Max output:
Input:
$
/ M UTF-8 bytes
Output:
$
/ M UTF-8 bytes
Fish Audio
text-to-speech
Fish-Speech-1.5
Release on: Nov 29, 2024
Total Context:
0K
Max output:
Input:
$
/ M UTF-8 bytes
Output:
$
/ M UTF-8 bytes

FunAudioLLM
text-to-speech
FunAudioLLM/CosyVoice2-0.5B
Release on: Dec 16, 2024
Total Context:
0K
Max output:
Input:
$
/ M UTF-8 bytes
Output:
$
/ M UTF-8 bytes
Model FAQs: Usage, Deployment
Learn how to use, fine-tune, and deploy this model with ease.
