- Home
- Speech to Text
Dashboard
How do you want to transcribe?
Free minutes are included. Upload a file or record audio to start.
Convert WAV to text on this page: choose a PCM WAV file, set the spoken language, and start a real transcription task in the workspace below. When the draft is ready, review the wording and timestamps, fix names and numbers, copy the passages you need, and download the edited transcript as TXT. Mono and stereo files are both accepted, and verified new accounts start with the current transcription-minute allowance, shown before you begin.
Dashboard
How do you want to transcribe?
Free minutes are included. Upload a file or record audio to start.
Verified accounts use the current Whisper AI transcription allowance. WAV audio is processed by the active cloud transcription provider, and finished tasks are saved in your recordings with the wav-to-text source tag.

This is the expected edited result for our own 30-second English narration, recorded as 16-bit PCM WAV in both mono and stereo. The speaker names herself, lists several numbers, and mentions the recorder settings. It is an illustrative example prepared without spending transcription credits, not a captured provider transcript. The timestamps show the editor view; a plain TXT export contains the text without them.
Source: original 30-second PCM WAV narration, mono and stereo (English, locally synthesized)
Edited TXT result
The sample is our own original narration, not a customer recording, and the transcript above is the expected edited text rather than a captured provider output. A machine draft can mishear names and numbers, so check the edited TXT against the audio before you publish or quote it.
WAV is an uncompressed container, so the file you upload is already the audio the recorder captured. Start from the original file and let the workspace produce a timestamped draft you can correct.
Upload a PCM WAV from this device, record audio in the browser, or import a supported public media URL. Mono and stereo files are both accepted, and 16-bit or 24-bit PCM is read as-is. When you know the spoken language, choose it before starting so the recognition pass is not guessing from a short sample.
Before a task is created, the workspace reads the WAV header, confirms the channel count, and applies the current file-size and minute limits. A file with a missing or damaged header, a non-PCM payload, more than two channels, or no audio data is rejected with a clear message instead of being sent for transcription.
Nothing is generated until you start the task, so opening the page never spends transcription minutes. When the draft is ready, replay the timestamps and scan for names, numbers, and passages affected by room tone or overlapping speech. The first result is a draft to correct, not a finished document.
Field recorders, audio editors, and studio exports often produce long WAV files with room tone, count-ins, or several takes. Plan the review around what the recording actually contains.
A handheld recorder often writes mono PCM at 44.1 or 48 kHz. Upload the original file rather than a re-encoded copy, choose the spoken language, and enable speaker separation only when the recording really contains more than one voice. Generated speaker labels group passages; they are not verified identities.
When you export a stem or a mixdown from a digital audio workstation, keep the sample rate and channel layout as they are. A stereo mix stays stereo and a mono narration stays mono; neither setting makes the transcript more accurate. Trim long silent leaders when you can, because they add duration without adding words.
A file renamed to .wav that actually contains MP3, AAC, or another codec fails the header check and is not sent for transcription. Keep the original recording and export a real PCM WAV, or switch to the matching tool for that format instead of forcing the extension.
The editor keeps your corrections with the task, so the file you download matches the version on screen instead of an older machine draft.
Correct names, numbers, acronyms, and specialist vocabulary directly in the transcript text. Replay the timestamp for any uncertain passage before changing it, and keep the spoken meaning rather than smoothing the speaker into different words.
Select and copy the passages you need for notes, an article, or a summary without leaving the page. Keep the timestamps if you plan to quote the recording, so a reviewer can find the original moment.
Download TXT for a plain, portable transcript, or use the workspace exports for DOCX, JSON, SRT, and VTT when the words need to become a document, structured data, or captions. Download after saving so the exported file matches the current version.
Comparison describes the checked pages at a high level. It does not repeat competitor pricing, ratings, certifications, or accuracy claims, and naming a product is not an endorsement.
| Decision factor | Whisper AI | Happy Scribe | Otter.ai |
|---|---|---|---|
| Starting from a WAV | Upload a PCM WAV from this page, record audio in the browser, or import a supported public media URL | Confirm the current upload, recording, and URL options on its live tool | Confirm the current upload, recording, and URL options on its live tool |
| Mono and stereo PCM handling | The WAV header is read before the task starts, so the workspace knows the channel count and sample rate and rejects damaged or data-free files | Confirm the current format and channel support on its live tool | Confirm the current format and channel support on its live tool |
| Editing before export | Fix wording, replay timestamps, copy passages, and keep edits saved with the task before downloading | An editor is offered; confirm which edits the current plan allows | An editor is offered; confirm which edits the current plan allows |
| Download formats | TXT download, plus DOCX, JSON, SRT, and VTT from the same transcript | Confirm the live export list and any plan gating | Confirm the live export list and any plan gating |
| Where the task lives | The finished task is saved in your recordings and tagged with its wav-to-text source | Confirm how the live product stores and reopens projects | Confirm how the live product stores and reopens projects |
| Trying the workflow | Verified new accounts start with the current transcription-minute allowance, shown before the task runs | Check the current free and paid plan terms | Check the current free and paid plan terms |
Whisper AI is the strongest fit when you want to start from a PCM WAV, confirm the channel layout before a task is created, see the transcript and its timestamps in an editor, keep the task with a traceable source, and download the edited text without switching to a separate workspace. Check each product's live terms before processing a long or confidential recording.
Yes. Mono and stereo PCM WAV files are both accepted. The workspace reads the header first, so it knows the channel count before the task starts. A stereo file is not more accurate than a mono file; the channel layout only describes how the recording was captured. If a file declares more than two channels, it is rejected so you can convert it first.
No, not for this workflow. WAV stores uncompressed audio, so there is no benefit in re-encoding it to MP3 or another lossy format just to upload it. Upload the PCM WAV directly within the current workspace file-size and minute limits shown before the task runs. Compress or split the file only when it is larger than the current limit, and keep the original recording.
Yes. When the draft is ready, open the transcript in the workspace, replay any timestamp, and correct names, numbers, and wording directly. Your edits are saved with the task, and the TXT or other export is built from the version you saved, so the download matches what you reviewed on screen.
The file is checked before a task is created. A missing or truncated header, a non-PCM payload, more than two channels, a file with no audio data, or a file over the current size or duration limit fails with a clear message, and no transcription task is started and no credits are spent. Keep the original recording and export a valid PCM WAV before trying again.
Common PCM WAV recordings at 44.1 kHz or 48 kHz with 16-bit or 24-bit samples are read as-is, in mono or stereo. The sample rate and bit depth describe the file; a higher sample rate does not make the transcript more accurate. Keep the sample rate your recorder or audio editor already uses instead of resampling just for upload.
Yes, once a task has been created. Signing in, returning from a recharge, or refreshing restores the task and its status instead of starting a second transcription, so you are not charged twice. If a refresh happens before the task is created and the browser no longer holds the local file, choose the file again to start.
This is a cloud transcription workflow, so the audio you submit is handled by the active transcription provider and stored so the task can be processed and reopened. Review the privacy policy before uploading confidential, medical, or client material, and avoid uploading recordings you are not authorized to process.
Written and checked by the Whisper AI Editorial Team · Last reviewed 2026-09-23 · Editorial policy