AI Audio & Voice
Text-to-speech, voice cloning and AI music generation. 32 tools tracked, updated daily.
Transformers
π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference a
Deer Flow
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message ga
GPT SoVITS
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
TTS
πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
ChatTTS
A generative speech model for daily dialogue.
VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
OpenVoice
Instant voice cloning by MIT and MyShell. Audio foundation model.
MockingBird
πClone a voice in 5 seconds to generate arbitrary speech in real-time
Index Tts
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
CosyVoice
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Dia
A TTS model capable of generating ultra-realistic dialogue in one pass.
Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognitio
Leon
π§ Leon is your open-source personal assistant.
Sherpa Onnx
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet conne
Supertonic
Lightning-Fast, On-Device, Multilingual TTS β running natively via ONNX.
Voice Pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processi
Edge Tts
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
Piper
A fast, local neural text to speech system
Amphion
Amphion (/Γ¦mΛfaΙͺΙn/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engin
Espnet
End-to-End Speech Processing Toolkit
EmotiVoice
EmotiVoice π: a Multi-Voice and Prompt-Controlled TTS Engine
VALL E X
An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/
Mlx Audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Sil
Vits
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
MeloTTS
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
Espeak Ng
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
SLAM LLM
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
ElevenLabs
The most realistic AI voice generator on the market
Suno
Create full songs from a text prompt
Udio
AI music studio for original tracks in seconds
Murf
Studio-quality AI voiceovers for professionals
PlayHT
AI voice API and ultra-realistic text to speech