About this tool
Audio to text uses speech recognition to produce editable text from a file and, where supported by the interface, live microphone capture. It creates a first draft for meetings, interviews, lessons, and voice notes but cannot guarantee verbatim accuracy or identify every speaker.
Accents, overlapping speakers, music, specialist terms, and poor microphones reduce accuracy. Proofread against the audio, especially names, numbers, negations, and accountable statements, and manually confirm timing and speakers in important records.
How to use it
Choose clear audio
Prefer a recording with forward speech, low noise, and stable volume.
Start transcription
Upload a file or authorize live capture and wait for model processing to complete.
Proofread and organize
Replay key locations and correct names, numbers, punctuation, and paragraph structure.
Supported range and limits
- Model
- Open-source Whisper running in-browser; the first run downloads tens of megabytes of model files
- Languages
- Multilingual, with solid results for English and Chinese; heavy accents and dense jargon reduce accuracy noticeably
- Time required
- Scales with audio length and device speed, and is usually well slower than real time; phones may not finish long material
- Output
- A plain-text transcript. Segmentation comes from the model, and there is no speaker separation
- Accuracy limits
- Background music, overlapping speakers and far-field recordings degrade it substantially; proofread the result
- Where it runs
- The model is downloaded, but inference on your audio happens in the browser and nothing is uploaded
When you would use it
Drafting meeting notes
Create searchable text quickly, then have participants verify decisions and action items.
Organizing an interview
Turn a long recording into searchable material for locating themes and quotations.
Drafting a transcript from an interview
Auto-transcribe first and proofread after; far quicker than typing from scratch while listening.
What to know before you start
- Automatic transcription must not be the sole legal, medical, or financial factual record.
- Live recognition may update in timed segments and is affected by network, device performance, and browser permissions.
- Obtain authorization and follow retention rules when handling other people’s voices or sensitive meetings.
Related concepts
- speech recognition
- The process of inferring words and text sequences from an acoustic signal.
- transcript
- A written record of audio content; an automatically generated version normally needs proofreading.