1. Key Features & Supported Languages
- Full List of 24 Languages: Arabic, Cantonese, Chinese, Czech, Dutch, English, Finnish, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Thai, Turkish, Ukrainian, Vietnamese.
- 21 Chinese Dialects: Anhui, Fujian, Gansu, Guizhou, Hebei, Henan, Hubei, Hunan, Jiangxi, Liaoning, Minnan, Ningxia, Shaanxi, Shandong, Shanghai, Shanxi, Sichuan, Tianjin, Wenzhou, Wu, Yunnan.
- Prompt-Based Voice Design: Generates novel timbres directly from natural language prompts without reference audio.
- Zero-Shot Voice Cloning: High-fidelity cloning with just a few seconds of audio.
- Free-Form Speech Editing: Enables word insertion, deletion, replacement, and fine-grained controls over pitch, pace, and volume.
⚠️ Notice on Multi-Speaker Feature: FireRedTTS3 officially omits the multi-speaker dialogue feature found in v2. Since v3 focuses on single-speaker precision, multi-role dialogues must now be handled at the application layer by generating sentences individually and stitching them together.
2. Key Differences from v2
- Architecture & Multi-Speaker Support: FireRedTTS2 uses a Dual-Transformer architecture optimized for multi-speaker conversational streaming. FireRedTTS3 adopts an LLM-DiT continuous generation architecture focusing on ultra-high cloning precision, voice design, and speech editing, officially omitting native multi-speaker support.
- Expanded Languages: v3 expands coverage to 24 languages and adds zero-shot cloning for 21 Chinese dialects.
3. Developer, Technology & Use Cases
- Developer: Created by Xiaohongshu's Super Intelligence (FireRed) Team.
- Technology: Built on an LLM-DiT frame with semantically enriched continuous speech representations to minimize cumulative generation errors.
- Use Cases: Ideal for AI voice acting, post-production speech modification, virtual character voice creation, and dialect audiobook narration.