Skip to content
AuK

AuK

Uses natural-language instructions to unify speech generation, editing, enhancement, and source separation

Download with PC Client

Features

Open SourceTTS

System Requirements

Minimum 32GB RAM. 36GB+ storage recommended.
macOS 15+: M-series chips required.
Windows 10/11: NVIDIA GPU with 12GB+ VRAM required.
Note: For NVIDIA GPUs, install a newer driver.

Introduction

AuK is an open-source speech generation and editing foundation model developed by Tencent. In simple terms, it can do much more than turn text into speech: it can also modify existing audio through natural-language instructions.

For example, you can provide a reference voice and ask AuK to speak new text in a similar voice. You can also ask it to replace a word in an existing recording, change the speaking speed or pitch, modify the speaker's emotion, remove noise, or separate different speakers from an audio recording.

One of AuK's key characteristics is that these different speech tasks are handled through a unified natural-language instruction interface, rather than requiring a completely separate model for each function.

Main Features

1. AI Speech Generation

AuK supports zero-shot text-to-speech (Zero-shot TTS). By providing a reference voice, users can generate new speech in a similar voice without specifically training a model for that speaker.

It also supports instruction-based TTS, where users can describe the desired voice characteristics with text without providing a reference recording.

2. Direct Editing of Existing Speech

AuK is not limited to text-to-speech generation. It can also edit existing recordings.

For example, it can:

  • replace, insert, or remove spoken words;
  • rewrite lyrics in a singing recording while preserving the melody and voice;
  • raise or lower pitch;
  • speed up or slow down speech;
  • adjust volume.

This makes speech editing more like editing a text document, allowing certain parts of a recording to be changed without having to recreate the entire recording.

3. Control Over Speaking Style and Voice Characteristics

AuK can modify how speech sounds while keeping the spoken content intact. It supports tasks such as:

  • changing emotion;
  • changing timbre;
  • removing regional accents;
  • converting between normal speech and whispering;
  • adding or removing non-verbal sounds such as breaths, laughter, and coughs.

In other words, AuK can control not only what is being said, but also how it is said.

4. Speech Enhancement and Audio Separation

For low-quality recordings, AuK can perform denoising, dereverberation, and speech restoration.

It can also separate voices from mixed audio—for example, keeping a particular speaker, separating speakers in a conversation, extracting vocals from music, or isolating a target speaker based on what they say.

Where Can It Be Used?

AuK can be viewed as an AI audio-editing engine, rather than simply another TTS system.

Potential applications include:

  • Content creation: narration, audiobooks, short videos, and podcasts;
  • Film and games: generating and modifying character dialogue, emotion, and voice characteristics;
  • Music production: lyric editing and vocal extraction;
  • Voice post-production: changing dialogue, pitch, or speaking speed without re-recording;
  • Audio restoration: reducing noise and reverberation and improving poor recordings;
  • Voice AI applications: providing speech generation and editing capabilities for other AI products.

A particularly important characteristic is that AuK brings speech generation, speech editing, enhancement, and source separation into a unified model and instruction framework.

Who Developed It?

AuK was developed and open-sourced by Tencent's Hunyuan team.

Both the source code and model weights have been publicly released. The project is licensed under the MIT License, allowing developers to use, modify, and build upon the released software and model components.

What Technology Does It Use?

AuK is a 1.5-billion-parameter speech foundation model, trained on millions of hours of diverse audio data.

Its implementation uses a modern generative AI architecture built around components including a text encoder, an audio VAE, and Transformer/flow-matching-based generation and editing modules. The source code specifically shows the use of Qwen2.5-Omni components, BigVGAN Flow VAE, and Flux2Edit/CFMEdit in the model pipeline.

The project provides two main model variants:

  • AuK: the base model, focused on high-quality speech generation;
  • AuK-Flash: a distilled version designed for faster inference, requiring only four inference steps.