Voices, voice models and supported languages
Updated September 15, 2026
The plugin offers 13 OpenAI voices and two voice models. You choose the model once in AI Text-to-Speech › Settings › General Settings › Voice Model, and pick a voice each time you generate audio. The default model is GPT-4o mini TTS (New).
Voice models
| GPT-4o mini TTS (New) | Basic TTS-1 (Old) | |
|---|---|---|
| API model name | gpt-4o-mini-tts | tts-1 |
| Voice styles and custom instructions | Yes | No |
| Cost used for estimates in the plugin | $0.012 per 1,000 characters | $0.015 per 1,000 characters |
| Maximum characters per API request | 6,000 | 2,000 |
The cost figures are what the plugin uses for its estimates. OpenAI sets and can change the real prices, so check the OpenAI pricing page and your usage dashboard for actual charges.
With TTS-1, the Style option isn’t shown in the post editor and no style instructions are sent to OpenAI.
Available voices
Marin, Cedar, Alloy, Ash, Ballad, Coral, Echo, Fable, Nova, Onyx, Sage, Shimmer and Verse. Marin, Cedar and Verse were added in version 3.2.0 and are labelled (NEW) in the voice list.
OpenAI doesn’t offer every voice on every model. Ballad, Verse, Marin and Cedar are designed for the newer GPT-4o mini TTS model, so if you use TTS-1, stick to Alloy, Ash, Coral, Echo, Fable, Nova, Onyx, Sage or Shimmer. If a voice isn’t supported, OpenAI returns an error and the plugin shows it in an alert.
You can hear samples of several voices on the AI Text to Speech plugin page. Every generation varies slightly, even with the same voice and text.
Set a default voice (PRO)
PRO users can choose which voice is pre-selected in AI Text-to-Speech › Settings › Automation › Default Voice. When an administrator generates audio, the voice and style they used are saved as the new defaults, so the last combination you used is remembered next time.
Supported languages
There’s no language setting. Write your post in the language you want, and OpenAI speaks it in that language. OpenAI supports over 50 languages, including English, Arabic, Chinese, Dutch, French, German, Greek, Hebrew, Hindi, Italian, Japanese, Korean, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, Ukrainian, Vietnamese and Welsh. The voices are tuned mainly for English, so results in other languages can vary.
One exception: currency amounts such as $25 or €3.50 are converted to English words before the text is sent. On non-English posts you may prefer to write amounts out in full, or use a [tts say=""] shortcode (PRO) to control how they’re read.


