Skip to content
FireRedTTS3

FireRedTTS3

An open-source LLM-DiT speech model by Xiaohongshu featuring zero-shot cloning for 24 languages and 21 dialects

Download with PC Client

Features

Open SourceTTS

Screenshots

FireRedTTS3 screenshot 1
FireRedTTS3 screenshot 2

System Requirements

16GB RAM recommended. 28GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11 64-bit: NVIDIA GPU with 8GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.

Introduction

1. Key Features & Supported Languages

  • Full List of 24 Languages: Arabic, Cantonese, Chinese, Czech, Dutch, English, Finnish, French, German, Greek, Hindi, Indonesian, Italian, Japanese, Korean, Polish, Portuguese, Romanian, Russian, Spanish, Thai, Turkish, Ukrainian, Vietnamese.
  • 21 Chinese Dialects: Anhui, Fujian, Gansu, Guizhou, Hebei, Henan, Hubei, Hunan, Jiangxi, Liaoning, Minnan, Ningxia, Shaanxi, Shandong, Shanghai, Shanxi, Sichuan, Tianjin, Wenzhou, Wu, Yunnan.
  • Prompt-Based Voice Design: Generates novel timbres directly from natural language prompts without reference audio.
  • Zero-Shot Voice Cloning: High-fidelity cloning with just a few seconds of audio.
  • Free-Form Speech Editing: Enables word insertion, deletion, replacement, and fine-grained controls over pitch, pace, and volume.

⚠️ Notice on Multi-Speaker Feature: FireRedTTS3 officially omits the multi-speaker dialogue feature found in v2. Since v3 focuses on single-speaker precision, multi-role dialogues must now be handled at the application layer by generating sentences individually and stitching them together.

2. Key Differences from v2

  • Architecture & Multi-Speaker Support: FireRedTTS2 uses a Dual-Transformer architecture optimized for multi-speaker conversational streaming. FireRedTTS3 adopts an LLM-DiT continuous generation architecture focusing on ultra-high cloning precision, voice design, and speech editing, officially omitting native multi-speaker support.
  • Expanded Languages: v3 expands coverage to 24 languages and adds zero-shot cloning for 21 Chinese dialects.

3. Developer, Technology & Use Cases

  • Developer: Created by Xiaohongshu's Super Intelligence (FireRed) Team.
  • Technology: Built on an LLM-DiT frame with semantically enriched continuous speech representations to minimize cumulative generation errors.
  • Use Cases: Ideal for AI voice acting, post-production speech modification, virtual character voice creation, and dialect audiobook narration.