toolgarden.xyz
中文

Text to speech

Turn Chinese or English text into natural Kokoro speech, adjust speed, preview it, and download a WAV file locally in your browser

Text to speech

High-quality Kokoro voices run locally in your browser. First use loads about 339 MB.0 / 2000 characters

Output audio

Enter text and generate speech to preview and download a WAV file here.

About this tool

Text to speech uses a Kokoro neural model in the browser to turn Chinese or English text into more natural and consistent speech. It no longer relies on operating-system voices, so results stay more predictable across devices.

Each language has two curated voices plus a speed control. Preview the result or download a standard WAV file for draft narration, pronunciation review, accessible listening, and voice-interface prototypes.

How to use it

  1. Enter text and choose a language

    Paste Chinese or English text, use punctuation for pauses, and keep each request within 2,000 characters.

  2. Choose a voice and speed

    Pick one of the two voices for the selected language and set a rate from 0.75x to 1.5x.

  3. Generate, preview, and download

    Let the browser synthesize the audio locally, listen to the result, then download the WAV file.

Supported range and limits

Synthesis engine
The bilingual Kokoro 82M model, using a high-precision ONNX build in the browser
Languages
Chinese and English
Voices
Chinese female, Chinese male, US female, and UK female
Adjustable
Speed from 0.75x to 1.5x
Output
24 kHz WAV audio with in-page preview and download
Where it runs
Text and generated audio stay in your browser and are not sent to ToolGarden servers

When you would use it

  • Drafting narration

    Turn a video script, product introduction, or lesson into downloadable reference audio.

  • Checking how content sounds

    Listen for awkward sentences, misplaced pauses, and incorrect number readings.

  • Accessible reading and prototypes

    Add stable Chinese or English speech to reading aids, interaction demos, and voice-interface prototypes.

What to know before you start

  • The first generation downloads about 339 MB of high-precision model data. Later visits can usually reuse the browser cache.
  • The model runs locally in the browser. Generation speed depends on your device, and longer text is synthesized in chunks.
  • Synthetic speech can still misread names, abbreviations, and unusual numbers, so review the complete output before publishing.
  • Do not use synthetic speech to impersonate a real person or create unauthorized deceptive material.

Related concepts

Kokoro
A lightweight open-weight text-to-speech model that turns text into natural spoken audio.
WAV
A common uncompressed PCM audio format suited to playback, editing, and further processing.

Frequently asked questions

Does text-to-speech upload my text?
No. A Kokoro model loaded in your browser generates the speech, and your text is not sent to ToolGarden servers. On first use, the browser downloads and caches the required model files.
What voices are available?
The language list is focused on Chinese and English, with two curated voices for each language instead of variable operating-system voice packs.
Can I export the speech to a file?
Yes. Preview the generated speech on the page, then download it as a standard WAV audio file.
Why is the first generation slower?
The browser must download and initialize about 339 MB of voice model data the first time. Once cached, later uses usually only wait for synthesis.
Can I save the generated speech?
Yes. Preview it in the output panel, then use the download button to save a WAV file.
Why are only Chinese and English available?
The current model is optimized for Chinese and English. A focused list ensures every choice has a stable voice instead of exposing inconsistent system voices.