Skip to content
MOSS-TTS 1.5

MOSS-TTS 1.5

An expressive open-source text-to-speech model supporting 31 languages, featuring stable zero-shot voice cloning and precise inline pause control

Features

Open SourceTTS

System Requirements

16GB RAM recommended. 23GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11 64-bit: NVIDIA GPU with 16GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.

Introduction

Only the Local Transformer v1.5 model is downloaded by default. Selecting the MOSS TTS v1.5 model will initiate an additional download of approximately 22GB of model files.

1. Product Positioning & Development Team MOSS-TTS-v1.5 is an open-source speech and sound generation model developed by MOSI.AI and the OpenMOSS team. As the latest iterative upgrade in the MOSS-TTS family, it is engineered for high-fidelity, high-expressiveness, and complex real-world applications, delivering an incredibly natural and controllable text-to-speech experience.

2. Core Features & Benefits (User-Friendly) For general users and content creators, MOSS-TTS-v1.5 offers several powerful and practical capabilities:

  • Zero-Shot Voice Cloning: By providing just a short snippet of reference audio, the model can instantly and accurately clone the speaker's voice profile, capturing their unique timber and prosody.
  • Explicit Pause Control: Users can inject inline pause markers, such as [pause 3.2s], directly into the text prompt. This allows the AI to pause precisely for the specified duration, making the speech rhythm sound human-like.
  • Robust Multilingual & Code-Switching Support: It supports 31 languages (including English, Chinese, Cantonese, Spanish, etc.). It effortlessly handles code-switching sentences (e.g., mixing English and Chinese) with smooth, natural transitions.
  • Stable Long-Form Generation: Designed to handle long articles or scripts without losing voice identity or degrading quality, while maintaining strict adherence to punctuation-driven intonations.

3. Application Scenarios This model is highly versatile and fits perfectly into various industries:

  • Content Creation & Audiobooks: Automating high-quality dubbing for novels, AI podcasts, and video narrations.
  • Gaming & Entertainment: Designing and tailoring expressive voices for virtual characters and NPC dialogue.
  • Smart Assistants & Customer Service: Enabling AI companions and customer service bots to interact with natural, empathetic, and fluent speech.

4. Underlying Technology Technically, MOSS-TTS-v1.5 is built on a Causal Transformer architecture. It leverages the advanced MOSS-Audio-Tokenizer to compress complex audio signals into efficient, unified semantic-acoustic representations. Utilizing an autoregressive generation pipeline with multi-scale Residual Vector Quantization (RVQ), the model guarantees high-quality audio decoding and superior long-context structural modeling.