Ming-omni-tts  by inclusionAI

Unified audio generation for speech, music, and sound with precise control

Created 8 months ago
266 stars

Top 96.9% on SourcePulse

GitHubView on GitHub
Project Summary

Ming-omni-tts offers a high-performance, unified approach to generating speech, music, and sound with precise user control. Targeting engineers, researchers, and power users, it provides efficient, controllable audio synthesis capabilities, including fine-grained vocal adjustments and zero-shot voice design. The project aims to deliver seamless, "in-the-scene" auditory experiences through its novel architecture and optimized inference.

How It Works

The core innovation lies in a custom 12.5Hz continuous VAE-based tokenizer that unifies speech, music, and general audio into a single latent space. This enables an end-to-end audio language model employing a unified LLM backbone, augmented with a Diffusion Head for enhanced audio quality. A patch-based generation strategy (patch size 4, look-back 32) combined with "Patch-by-Patch" compression achieves a highly efficient 3.1Hz inference frame rate. This architecture balances local acoustic detail with long-range structural coherence, leading to competitive performance across various audio synthesis tasks.

Quick Start & Requirements

Installation can be done via pip (pip install -r requirements.txt) or Docker (docker pull yongjielv/ming_uniaudio:v1.1). Model downloads are available via ModelScope (modelscope download ...) or Hugging Face. The project requires GPU acceleration (tested on NVIDIA H800-80GB/H20-96G with CUDA 12.4). Example usage is provided in cookbooks/cookbook.ipynb. Further details can be found on the project's blog: https://xqacmer.github.io/Ming-Flash-Omni-V2-TTS/.

Highlighted Details

  • Fine-grained Vocal Control: Supports precise adjustments for speech rate, pitch, volume, emotion (46.7% accuracy), and dialect (93% accuracy for Cantonese).
  • Intelligent Voice Design: Offers 100+ built-in voices and zero-shot voice design via natural language descriptions, matching Qwen3-TTS performance on Instruct-TTS-Eval-zh.
  • Unified Generation: The first autoregressive model capable of jointly generating speech, ambient sound, and music within a single channel.
  • High-efficiency Inference: Achieves 3.1Hz inference speed through "Patch-by-Patch" compression, enabling low-latency podcast-style audio.
  • Professional Text Normalization: Accurately narrates complex mathematical and chemical expressions with a CER of 1.97%, comparable to Gemini-2.5 Pro.
  • Performance Benchmarks: Outperforms CosyVoice3 in dialect and emotion control, shows competitive zero-shot TTS results, and excels in Text-to-BGM and Text-To-Audio tasks.

Maintenance & Community

The provided README does not detail specific maintenance contributors, sponsorships, or community channels like Discord or Slack. The project blog serves as an update resource.

Licensing & Compatibility

No specific open-source license is mentioned in the provided README. This lack of explicit licensing information presents a significant caveat for commercial use or integration into closed-source projects.

Limitations & Caveats

The README does not explicitly list limitations such as alpha status or known bugs. However, the model download process can be time-consuming (minutes to hours). Testing was performed on high-end NVIDIA GPUs with CUDA 12.4, suggesting potential hardware dependencies or performance variations on different systems. The absence of clear licensing information is a critical adoption blocker.

Health Check
Last Commit

7 months ago

Responsiveness

Inactive

Pull Requests (30d)
0
Issues (30d)
1
Star History
0 stars in the last 30 days

Explore Similar Projects

Starred by Christian Laforte Christian Laforte(Distinguished Engineer at NVIDIA; Former CTO at Stability AI), Chip Huyen Chip Huyen(Author of "AI Engineering", "Designing Machine Learning Systems"), and
1 more.

Amphion by open-mmlab

0%
10k
Toolkit for audio, music, and speech generation research
Created 2 years ago
Updated 6 months ago
Starred by Chip Huyen Chip Huyen(Author of "AI Engineering", "Designing Machine Learning Systems"), Jeff Hammerbacher Jeff Hammerbacher(Cofounder of Cloudera), and
2 more.

AudioGPT by AIGC-Audio

0%
10k
Audio processing and generation research project
Created 3 years ago
Updated 2 years ago
Feedback? Help us improve.