Open-weights emotive TTS · live demo

Miso One — Emotive AI Text to Speech

The MisoTTS model by Miso Labs, running live in your browser

Miso One is an open-weights, ~8B-parameter text-to-speech model built to sound genuinely human — warm, expressive and emotive. Type any English text below and hear lifelike AI speech, or clone a voice from a short clip.

~8B parametersOne-shot voice cloningOpen weights (modified MIT)EnglishRuns in your browser
Miso One · MisoTTS live demo
Open in Hugging Face

The demo is the official MisoTTS Space hosted on Hugging Face. The first generation may take a little longer while the GPU warms up. Miso One and MisoTTS are products of Miso Labs; this page is an independent guide with the official demo embedded.

Overview

What is Miso One?

Miso One is a highly emotive AI text-to-speech model from Miso Labs, a Y Combinator-backed voice-AI startup. Released in early June 2026 under the technical name MisoTTS, it is an ~8-billion-parameter speech synthesis model designed to make AI voices sound genuinely human — with real warmth, pacing and emotion where most text-to-speech still falls flat.

Under the hood, Miso One pairs a Llama 3.2-style backbone with a smaller autoregressive audio decoder — an RVQ Transformer inspired by Sesame's CSM architecture — and generates Mimi audio codes from your text and optional audio context. It supports one-shot voice cloning from a roughly 10-second clip, and Miso Labs published it with open weights so developers can run it themselves.

The interactive panel at the top of this page is the official MisoTTS Hugging Face Space, so you can try Miso One as an AI voice generator right now — no install required.

Features

Why Miso One stands out

Emotive delivery, voice cloning and open weights make Miso One a capable, developer-friendly text-to-speech model.

Genuinely emotive speech

Miso One is built to read like a real person — with warmth, pacing and emotion — instead of the flat, robotic delivery most text-to-speech still produces.

One-shot voice cloning

Clone a voice from roughly a 10-second reference clip, then have Miso One narrate new lines in that voice. Use cloning only with the speaker's consent.

Low-latency generation

Miso Labs reports first-token latency around 110 ms for responsive, conversational playback. Treat the number as a vendor benchmark.

~8B-parameter architecture

A Llama 3.2-style backbone paired with a smaller audio decoder — an RVQ Transformer inspired by Sesame's CSM — generating Mimi audio codes from text.

Open weights, self-hostable

Released with open weights under a modified MIT license, so you can run Miso One locally and keep your audio and text on your own machine.

Built for developers

Pull the MisoLabs/MisoTTS weights from Hugging Face and integrate them into your own stack. Miso Labs lists a hosted API as coming soon.

Expressive English voices

Miso One focuses on natural English speech today, with lifelike intonation, emphasis and phrasing that follow your punctuation.

Tone-aware output

The model conditions on both your text and optional audio context, so the delivery can match the mood and energy you are aiming for.

How to use

Generate speech with Miso One in four steps

Everything happens in the embedded demo above — no account or setup needed to start.

1

Type or paste your text

Enter the English text you want spoken into the Miso One demo above. Punctuation and natural phrasing guide the emotion and pacing.

2

Add a reference voice (optional)

Provide a short, clean reference clip — around 10 seconds — if you want Miso One to clone a specific voice for the output.

3

Adjust and generate

Use any expressiveness or length controls the demo exposes, then generate. The first run may be slower while the Hugging Face Space warms up.

4

Play, download or self-host

Listen to and download the generated audio. For production use, run the open MisoTTS weights yourself so your data never leaves your infrastructure.

Use cases

What you can build with Miso One

Emotive AI speech fits a wide range of voice and audio products.

Voiceover & narration

Produce expressive narration for explainers, ads and social video without booking a studio session.

Conversational voice agents

Give assistants, IVRs and support bots a voice that sounds human, thanks to low-latency, emotive delivery.

Audiobooks & e-learning

Turn long-form scripts and courses into natural-sounding speech with consistent pacing and tone.

Game & character dialogue

Prototype or ship character lines with distinct, emotional voices using one-shot voice cloning.

Accessibility

Read articles, documents and UI aloud in a voice that is easier and more pleasant to listen to.

Podcasts & content creation

Draft intros, ad reads and segments, or generate a consistent voice for repeatable content.

Tips

Tips for better results

  • Write the way you want it spoken — commas, periods and line breaks shape Miso One's pacing and emphasis.
  • For voice cloning, use a clean, roughly 10-second mono clip with minimal background noise for the best match.
  • Split long scripts into natural sentences; generate section by section rather than one huge block.
  • The first generation can be slow while the ZeroGPU Space spins up — later runs are faster.
  • Don't paste sensitive or private text into the public demo. For confidential work, self-host the open weights.
FAQ

Miso One frequently asked questions

Common questions about Miso One, MisoTTS and Miso Labs.

What is Miso One?

Miso One is an open-weights, highly emotive AI text-to-speech (TTS) model from Miso Labs. It generates lifelike, expressive English speech from text and can clone a voice from a short reference clip. The interactive panel above is the official MisoTTS demo running on Hugging Face.

Is Miso One the same as MisoTTS?

Yes. “Miso One” is the product name and “MisoTTS” is the open-source release and repository name for the same model. You'll see the MisoTTS name on the GitHub project and the Hugging Face model and Space.

Who created Miso One?

Miso One was built by Miso Labs, a Y Combinator-backed startup focused on emotive voice AI. It was released with open weights in early June 2026.

Is Miso One free and open source?

The demo above is free to try in your browser, and Miso Labs published the model with open weights under a modified MIT license. Review the exact license terms in the official repository before using it commercially.

What languages does Miso One support?

Miso One is focused on English today. Other languages are not documented as supported at launch.

Can Miso One clone a voice?

Yes. Miso Labs describes one-shot voice cloning from roughly a 10-second reference clip. Only clone voices you have permission to use, and follow applicable laws and platform rules.

How large is the model and how fast is it?

Miso One is described as an ~8-billion-parameter model (a Llama 3.2-style backbone plus an audio decoder). Miso Labs reports first-token latency around 110 ms; that figure is a vendor claim rather than an independent benchmark.

Does my data stay private?

Because the weights are open, you can self-host Miso One and keep audio and text entirely on your own machine. The hosted demo above runs on Hugging Face, so avoid pasting confidential content into it.

Hear Miso One for yourself

Type a line, pick a voice and listen to emotive AI speech in seconds — then explore the open weights to build with MisoTTS.