AI voice & speech models · directory

AI Models

Explore the AI voice and speech models we host and document. Try interactive demos, then dive into a dedicated guide for each one. New models are added here as they launch.

The collection

Available AI models

Each model has its own page with a live demo, capabilities and an honest, hands-on guide.

Speech Recognition

Nemotron 3.5 ASR

NVIDIA's open 0.6B multilingual streaming speech-recognition model. Try Nemotron AI audio-to-text live — upload a clip or paste a URL and get punctuated transcripts.

Streaming ASR~40 localesOpen weightsAudio to text
Explore Nemotron 3.5 ASR
Speech Recognition Alternative

Muse Voice Transcribe Alternative

An independent workflow for evaluating Muse Voice Transcribe use cases. It currently uses ElevenLabs Scribe v2 with an eligible Whisper fallback, not the Meta Model API.

Speech to textSpeaker labelsWord timingAlternative
Explore Muse Voice Transcribe Alternative
Text to Audio

Seed Audio 1.0

ByteDance's Seed Audio 1.0 model on fal can generate audio from a prompt, reference audio or an image. Configure the API request and review responsible-use notes.

Text to audioReference audioImage guidancefal API
Explore Seed Audio 1.0
Text to Speech

Miso One

Miso Labs' open-weights, highly emotive AI text-to-speech model (MisoTTS). Try the live demo, clone a voice from a short clip and hear lifelike speech.

Emotive TTSVoice cloningOpen weightsEnglish
Explore Miso One
Multilingual Text to Speech

OmniVoice

k2-fsa's multilingual zero-shot TTS model for 600+ languages. Try the live Hugging Face demo for voice cloning, voice design and natural AI speech.

600+ languagesVoice cloningVoice designApache-2.0
Explore OmniVoice
Multilingual Text to Speech

VoxCPM

OpenBMB's tokenizer-free VoxCPM2 TTS model for 30 languages, voice design, controllable voice cloning and 48 kHz speech. Try the official Hugging Face demo.

30 languagesVoice designVoice cloning48 kHz
Explore VoxCPM

More models coming soon

We're adding more AI voice and speech models to this directory. Check back, or start with Nemotron 3.5 ASR, Seed Audio 1.0, Miso One, OmniVoice or VoxCPM today.