An expressive open-source text-to-speech model supporting 31 languages, featuring stable zero-shot voice cloning and precise inline pause control
16GB RAM recommended. 23GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11 64-bit: NVIDIA GPU with 16GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.Only the Local Transformer v1.5 model is downloaded by default. Selecting the MOSS TTS v1.5 model will initiate an additional download of approximately 22GB of model files.
1. Product Positioning & Development Team MOSS-TTS-v1.5 is an open-source speech and sound generation model developed by MOSI.AI and the OpenMOSS team. As the latest iterative upgrade in the MOSS-TTS family, it is engineered for high-fidelity, high-expressiveness, and complex real-world applications, delivering an incredibly natural and controllable text-to-speech experience.
2. Core Features & Benefits (User-Friendly) For general users and content creators, MOSS-TTS-v1.5 offers several powerful and practical capabilities:
[pause 3.2s], directly into the text prompt. This allows the AI to pause precisely for the specified duration, making the speech rhythm sound human-like.3. Application Scenarios This model is highly versatile and fits perfectly into various industries:
4. Underlying Technology Technically, MOSS-TTS-v1.5 is built on a Causal Transformer architecture. It leverages the advanced MOSS-Audio-Tokenizer to compress complex audio signals into efficient, unified semantic-acoustic representations. Utilizing an autoregressive generation pipeline with multi-scale Residual Vector Quantization (RVQ), the model guarantees high-quality audio decoding and superior long-context structural modeling.