Try the workflow now
Evaluate upload, recording, transcript editing, speaker labels, timestamps, and exports without waiting for another provider integration.
Upload audio or video, or record a short clip. Get an editable multilingual transcript with optional speaker labels, word-level timing, and important-word prompting.

Choose one file to continue in the transcription workspace.
Verified accounts use the current Whisper AI transcription allowance. Media is sent to the active cloud transcription provider.
Muse Voice Transcribe is Meta's speech-to-text model for streaming recognition, speech endpoint detection, multilingual conversations, and optional speaker diarization. Meta announced it on September 1, 2026 and documented both a recording endpoint and a real-time WebSocket endpoint through the Meta Model API.
Those capabilities describe Meta's product. The interactive Whisper AI tool above currently uses a different provider, disclosed before submission.
Test the core batch-transcription workflow today while keeping the model and provider boundary clear.
Evaluate upload, recording, transcript editing, speaker labels, timestamps, and exports without waiting for another provider integration.
Scribe v2 supports multilingual transcription, mixed-language handling, diarization, word timing, and keyterm prompting.
The page describes a comparable workflow, not identical model output, and shows whether Scribe or the Whisper fallback processed the task.
Usage, checkout attribution, purchases, and official-engine requests show whether funding the Meta integration is justified.
Start with a file or a short browser recording, then shape the transcript for review and export.
Upload a supported file or record a clip.
Choose Auto or a spoken-language hint.
Enable speaker labels for conversations.
Add important names or technical terms.
Review the transcript and export the format you need.
A factual comparison of the documented Meta product and the provider currently powering this validation page.
| Area | Meta Muse Voice Transcribe | Current Whisper AI validation tool |
|---|---|---|
| Active provider | Meta Model API | ElevenLabs Scribe v2; Replicate fallback |
| Uploaded files | Official recording endpoint | Scribe v2 batch transcription |
| Real-time streaming | Separate Meta WebSocket endpoint | Not enabled on this page |
| Languages | Meta says 70+ trained, 25 extensively verified at launch | Scribe v2 documents 90+ languages |
| Mixed-language audio | Code switching is a launch feature | Scribe v2 smart multi-language transcription |
| Speaker labels | Native diarization mode | Native Scribe diarization, up to 32 speakers through the API |
| Important terms | Keywords, language, and context bias | Scribe keyterms and language code |
| Input | Public file recipe documents WAV constraints | Major audio and video formats accepted by Scribe |
No. The first release processes an uploaded file or a browser recording after you stop. Meta's Muse product and ElevenLabs each document separate real-time APIs, but neither live integration is enabled on this page.
Provider, model, streaming, diarization, and fallback details for this validation release.
Muse Voice Transcribe is Meta's speech-to-text model for streaming recognition, endpoint detection, multilingual speech, and speaker diarization. Meta announced recording and real-time API workflows in September 2026.
No. The current validation tool uses ElevenLabs Scribe v2, with Replicate Whisper available as an operational fallback. The active provider is disclosed before submission and on completed validation tasks.
It offers a comparable batch workflow—multilingual speech to text, speaker labels, timestamps, and important-word hints—without claiming to run Meta's model. Model output and accuracy can differ.
Yes. Meta documents an HTTP endpoint for completed recordings and a WebSocket endpoint for live audio. Whisper AI is measuring demand before deciding whether to fund an official integration.
No. Uploads and browser recordings are processed after submission; this page does not show partial text while you speak.
Scribe v2 supports optional speaker diarization and automatic multi-language transcription. Results still vary with noise, accents, overlapping speech, and recording quality.
For eligible temporary provider failures, Whisper AI may use its existing Replicate Whisper fallback without charging the transcription twice. The completed task identifies the provider used.
No. Whisper AI is an independent product and is not affiliated with or endorsed by Meta.
Register demand without joining an email list, or continue with the disclosed alternative.
This records product interest; it does not subscribe you to email.