AI Models
Explore the AI voice and speech models we host and document. Try interactive demos, then dive into a dedicated guide for each one. New models are added here as they launch.
Available AI models
Each model has its own page with a live demo, capabilities and an honest, hands-on guide.
Nemotron 3.5 ASR
NVIDIA's open 0.6B multilingual streaming speech-recognition model. Try Nemotron AI audio-to-text live — upload a clip or paste a URL and get punctuated transcripts.
Seed Audio 1.0
ByteDance's Seed Audio 1.0 model on fal can generate audio from a prompt, reference audio or an image. Configure the API request and review responsible-use notes.
Miso One
Miso Labs' open-weights, highly emotive AI text-to-speech model (MisoTTS). Try the live demo, clone a voice from a short clip and hear lifelike speech.
OmniVoice
k2-fsa's multilingual zero-shot TTS model for 600+ languages. Try the live Hugging Face demo for voice cloning, voice design and natural AI speech.
VoxCPM
OpenBMB's tokenizer-free VoxCPM2 TTS model for 30 languages, voice design, controllable voice cloning and 48 kHz speech. Try the official Hugging Face demo.
More models coming soon
We're adding more AI voice and speech models to this directory. Check back, or start with Nemotron 3.5 ASR, Seed Audio 1.0, Miso One, OmniVoice or VoxCPM today.
Looking for transcription?
These models pair well with our core speech-to-text workspace.
