AI Audio & Voice

Text-to-speech, voice cloning and AI music generation. 32 tools tracked, updated daily.

Transformers

πŸ€— Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference a

freemiumβ˜… 166.4k

Deer Flow

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message ga

freeβ˜… 82.7k

GPT SoVITS

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

freemiumβ˜… 61.9k

TTS

πŸΈπŸ’¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

freemiumβ˜… 46.0k

ChatTTS

A generative speech model for daily dialogue.

freemiumβ˜… 39.9k

VoxCPM

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

freemiumβ˜… 37.8k

OpenVoice

Instant voice cloning by MIT and MyShell. Audio foundation model.

freemiumβ˜… 37.6k

MockingBird

πŸš€Clone a voice in 5 seconds to generate arbitrary speech in real-time

freemiumβ˜… 36.9k

Index Tts

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

freemiumβ˜… 24.1k

CosyVoice

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

freemiumβ˜… 23.7k

Dia

A TTS model capable of generating ultra-realistic dialogue in one pass.

freemiumβ˜… 19.4k

Speech

A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognitio

freemiumβ˜… 18.5k

Leon

🧠 Leon is your open-source personal assistant.

freeβ˜… 17.5k

Sherpa Onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet conne

freemiumβ˜… 14.9k

Supertonic

Lightning-Fast, On-Device, Multilingual TTS β€” running natively via ONNX.

freemiumβ˜… 13.8k

Voice Pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processi

freemiumβ˜… 12.8k

Edge Tts

Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key

freemiumβ˜… 12.0k

Piper

A fast, local neural text to speech system

freemiumβ˜… 11.3k

Amphion

Amphion (/Γ¦mˈfaΙͺΙ™n/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engin

freemiumβ˜… 10.3k

Espnet

End-to-End Speech Processing Toolkit

freemiumβ˜… 10.0k

EmotiVoice

EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine

freemiumβ˜… 8.5k

VALL E X

An open source implementation of Microsoft's VALL-E X zero-shot TTS model. Demo is available in https://plachtaa.github.io/vallex/

freeβ˜… 7.9k

Mlx Audio

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Sil

freemiumβ˜… 7.9k

Vits

VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech

freemiumβ˜… 7.9k

MeloTTS

High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.

freemiumβ˜… 7.6k

Espeak Ng

eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.

freeβ˜… 6.9k

SLAM LLM

A Framework for Speech, Language, Audio, Music Processing with Large Language Model

freemiumβ˜… 1.1k

ElevenLabs

The most realistic AI voice generator on the market

freemium

Suno

Create full songs from a text prompt

freemium

Udio

AI music studio for original tracks in seconds

freemium

Murf

Studio-quality AI voiceovers for professionals

paid

PlayHT

AI voice API and ultra-realistic text to speech

paid