toolgarden.xyz
中文

Audio to text

Free online audio transcription tool that loads an open-source Whisper model in the browser and converts speech to text

Upload audio

Use the microphone for near real-time transcription. Audio stays in this browser session.

Transcript

About this tool

Audio to text uses speech recognition to produce editable text from a file and, where supported by the interface, live microphone capture. It creates a first draft for meetings, interviews, lessons, and voice notes but cannot guarantee verbatim accuracy or identify every speaker.

Accents, overlapping speakers, music, specialist terms, and poor microphones reduce accuracy. Proofread against the audio, especially names, numbers, negations, and accountable statements, and manually confirm timing and speakers in important records.

How to use it

  1. Choose clear audio

    Prefer a recording with forward speech, low noise, and stable volume.

  2. Start transcription

    Upload a file or authorize live capture and wait for model processing to complete.

  3. Proofread and organize

    Replay key locations and correct names, numbers, punctuation, and paragraph structure.

Supported range and limits

Model
Open-source Whisper running in-browser; the first run downloads tens of megabytes of model files
Languages
Multilingual, with solid results for English and Chinese; heavy accents and dense jargon reduce accuracy noticeably
Time required
Scales with audio length and device speed, and is usually well slower than real time; phones may not finish long material
Output
A plain-text transcript. Segmentation comes from the model, and there is no speaker separation
Accuracy limits
Background music, overlapping speakers and far-field recordings degrade it substantially; proofread the result
Where it runs
The model is downloaded, but inference on your audio happens in the browser and nothing is uploaded

When you would use it

  • Drafting meeting notes

    Create searchable text quickly, then have participants verify decisions and action items.

  • Organizing an interview

    Turn a long recording into searchable material for locating themes and quotations.

  • Drafting a transcript from an interview

    Auto-transcribe first and proofread after; far quicker than typing from scratch while listening.

What to know before you start

  • Automatic transcription must not be the sole legal, medical, or financial factual record.
  • Live recognition may update in timed segments and is affected by network, device performance, and browser permissions.
  • Obtain authorization and follow retention rules when handling other people’s voices or sensitive meetings.

Related concepts

speech recognition
The process of inferring words and text sequences from an acoustic signal.
transcript
A written record of audio content; an automatically generated version normally needs proofreading.

Frequently asked questions

Does audio-to-text upload my recording?
No. Transcription uses an open-source Whisper model loaded locally in your browser, so recordings are never uploaded.
Does it support Chinese and other languages?
Yes. The Whisper model recognises many languages including Chinese and English, and handles mixed-language audio.
Why is the first run slower?
The first transcription downloads the Whisper model in your browser; it is then cached and reused, so later runs are much faster.
Is my audio uploaded?
No. The model files are downloaded to the browser, but inference happens locally and the audio never leaves your device. That is what makes it usable for interviews, medical notes and other sensitive recordings.
Why is it so slow?
Inference runs on your device and speed follows your CPU. It is usually well slower than real time; ten minutes of audio can take several minutes to process, and phones may not finish at all.
How accurate is it?
Very good on a single clear speaker, but background music, overlapping voices, far-field recording and specialist terminology all pull it down sharply. Treat the output as a draft to proofread, not a finished transcript.