- Home
- Speech to Text
Dashboard
How do you want to transcribe?
Free minutes are included. Upload a file or record audio to start.
Convert a video to text on this page: choose an MP4, MOV, or WEBM file, set the spoken language, and start a real transcription task in the workspace below. When the draft is ready, follow the timestamps to the exact moment, fix names and numbers, copy the passages you need, and download the words from your video as TXT or another export format. Verified new accounts start with the current transcription-minute allowance, shown before you begin.
Dashboard
How do you want to transcribe?
Free minutes are included. Upload a file or record audio to start.
Verified accounts use the current Whisper AI transcription allowance. The video's audio track is processed by the active cloud transcription provider, and finished tasks are saved in your recordings with the video-to-text source tag.

This is the expected edited result for two of our own English clips: a 41-second screen-recording WEBM that walks through a dashboard, and a 22-second phone MOV recorded straight after it. The timestamps show where each cue sits, and the same cue times carry into the subtitle export. It is an illustrative example prepared without spending transcription credits, not a captured provider transcript.
Sources: 41-second screen-recording WEBM (desktop walkthrough) and 22-second phone MOV (spoken note), both English
Edited timestamped transcript
The two clips are our own recordings, and the transcript above is the expected edited text rather than a captured provider output. A machine draft can mishear names, numbers, and product terms, so check the edited text against the video before you publish or quote it. A silent or music-only clip has no speech to place on a timeline, so the result comes back with no words instead of filling the transcript with invented words; check that the audio track really contains speech before you rely on the output.
Start from the video file itself. The workspace reads the audio track, so you do not have to extract or convert the sound first.
Upload an MP4, MOV, or WEBM from this device, or import another supported media source from the workspace. One video per task keeps the timestamps and the transcript easy to check. If a clip is long, a short representative section is enough to confirm that the language, names, and audio quality transcribe well.
Choose the spoken language before you start when you know it. Setting it up front helps short clips and noisy recordings, because the recognition pass is not guessing the language from a small sample. Leave it on automatic only when you genuinely do not know what is being said.
The workspace checks the video, states the current file-size and duration limits before the task runs, and shows the credit cost against your balance. Nothing is generated until you start the task, so opening the page never spends transcription minutes. A silent clip returns no words rather than guessed text, and an unsupported codec or damaged container stops with a readable error.
A timestamp is the bridge between what you hear and where it happens in the video. Keep the draft and the source side by side so every correction can be checked at its moment.
Use the timestamps to replay the passage you are unsure about. Confirm names, numbers, product terms, and anything spoken over music or background noise before you accept it. The first result is a draft to verify, not a finished document.
Correct names, numbers, acronyms, and specialist vocabulary directly in the transcript text. Your edits are saved with the task, so the TXT, DOCX, and JSON exports match the version on screen instead of an older machine draft. Keep the spoken meaning rather than smoothing the speaker into different words.
Copy the passages you need for notes, an article, or a summary, and keep the timestamps if a reviewer will look up the original moment. Download TXT for a plain transcript, or export DOCX, JSON, SRT, or VTT when the words need to become a document, structured data, or timed captions.
The file you have, where it lives, and what you will do with the words decide which workflow and export fit best.
This page defaults to a video file on your device, which is the fastest path when you control the source. For a public platform page such as YouTube or Vimeo, use the workspace's Import a media URL tab and paste the page link; the resolver must approve the public URL before a task can start. If you already downloaded the video, uploading the file here avoids that extra step.
Choose a transcript when the words are the deliverable, and captions when the timing is. Create the transcript first for the words, then open the subtitle view and correct each cue's text and timing before you export SRT or VTT. In the workspace, full-text edits change the readable transcript copy only; subtitle files use the cue list, so edit the cues themselves, then review line length, reading speed, and cue breaks in the player that will actually show them.
If you only need a playable audio file, use a format converter instead of spending transcription minutes. If your source is a single MP4 and you want the same timestamped output, the MP4-to-text page covers that specific format. If you want to test the workflow first, start with the verified-signup allowance on this page or the free audio-to-text entry.
Comparison describes the checked pages at a high level. It does not repeat competitor pricing, ratings, certifications, or accuracy claims, and naming a product is not an endorsement.
| Decision factor | Whisper AI | Happy Scribe | VEED |
|---|---|---|---|
| Starting from a video | Upload an MP4, MOV, or WEBM on this page, or import a supported public media URL from the workspace | Confirm the current upload and URL options on its live tool | Confirm the current upload and URL options on its live tool |
| Timestamps for navigation | Clickable timestamps in the editor, and the same cue times feed the SRT and VTT exports | Timed text is offered; confirm how the live editor exposes each cue | Timed captions live on a video timeline; confirm how the live tool handles them |
| Editing before export | Fix wording, replay a timestamp, and keep edits saved with the task before downloading | An editor is offered; confirm which edits the current plan allows | Editing happens on the video timeline; confirm the current plan gating |
| Where the task lives | The finished task is saved in your recordings and tagged with its video-to-text source | Confirm how the live product stores and reopens projects | Confirm how the live product stores and reopens projects |
| Trying the workflow | Verified new accounts start with the current transcription-minute allowance, shown before the task runs | Check the current free and paid plan terms | Check the current free and paid plan terms |
Whisper AI is the strongest fit when you want to start from a video file, read a timestamped transcript, correct it in place, and export the words or captions without leaving the page. A reviewer who needs to compare products should check each product's live terms before processing a long or confidential recording.
This page is built around MP4, MOV, and WEBM, which are the video formats the workspace handles by default, and the upload control also accepts other common containers such as MKV and M4V. The workspace checks the actual media rather than trusting the file name, so a file that is renamed to a supported extension but contains unreadable or unsupported video fails with a clear message instead of producing a junk transcript.
Yes. Upload the video and the workspace reads its audio track directly, so you do not need to extract a separate MP3 or WAV first. That keeps the source and the transcript in one task, which is useful because a corrected transcript and its timestamps stay tied to the exact video you uploaded.
This page defaults to a video file on your device. For a public platform page such as YouTube or Vimeo, open the workspace's Import a media URL tab and paste the page link; the resolver must approve the public URL before a task starts. If you already have the video file, uploading it here is the more direct path and avoids depending on a platform page staying reachable.
A silent or music-only clip has no spoken words to place on a timeline, so the transcript comes back with no words rather than a page of guessed lines; it does not invent speech that is not there. Check that the audio track really contains speech before you rely on the output, and treat an empty transcript as a sign that the recording, not the recognizer, needs attention.
Yes. When the draft is ready, open the transcript in the workspace, replay any timestamp, and correct names, numbers, and wording directly. Your edits are saved with the task, so the TXT, DOCX, and JSON exports are built from the version you saved and the download matches what you reviewed on screen.
The cue times come from the same transcription task, so an SRT or VTT export uses the same moments you hear in the video. Editing the full transcript text updates the readable copy only; subtitle files are built from the cue list in the subtitle view, so correct each cue's text and timing there before you export SRT or VTT, then review line length and reading speed in the player that will display the captions.
Yes, once a task has been created. Signing in, returning from a recharge, or refreshing restores the task and its status instead of starting a second transcription, so you are not charged twice. If a refresh happens before the task is created and the browser no longer holds the local file, choose the video again to start.
Written and checked by the Whisper AI Editorial Team · Last reviewed 2026-09-23 · Editorial policy