Skip to main content

Generating audio with text-to-speech

Turn written text into a narrated audio file, choose a voice, and save it to your media library.

Written by Cody Iddings

Text-to-speech turns text you type into a narrated audio file you can use anywhere audio is used — an audio activity, an intro, an expert tip. It's generated with AI voices, so you don't need a microphone or a recording session.

Text-to-speech is in beta, and it's switched on per organisation. If you can see a Text-to-speech tab in the media library, you have it.

Generating audio

  1. Open your media library, or the media picker inside an experience.

  2. Switch to the Text-to-speech tab.

  3. Type or paste your script into the text box. There's a 5,000 character limit, and the count is shown as you type.

  4. Pick a voice.

  5. Click Generate.

When it's ready, a player appears at the bottom so you can listen before you commit to anything.

Choosing a voice

Voices are grouped by accent, each one listed with its name and whether it reads as male or female. You can play a sample of any voice before choosing it, so you don't have to generate your script to find out how a voice sounds.

Fine-tuning the delivery

Under Settings you can shape how the voice performs your text:

  • Speed — how fast the voice reads. The usable range is fairly narrow; pushing it to the ends starts to sound unnatural.

  • Stability — lower is more expressive and more variable between takes, higher is flatter and more consistent.

  • Similarity — how closely the output sticks to the original voice.

  • Style exaggeration — how much the voice leans into the emotion it reads in your text.

  • Speaker boost — sharpens the resemblance to the source voice.

Double-click any slider to put it back to its default.

If you don't touch these at all you'll get a sensible, natural read. They're worth reaching for when a specific line isn't landing — not something to set before your first attempt.

Naming and saving

Once you've generated something, give it an audio name — this is what it will be called in your media library, so it's worth being specific. A name is suggested for you automatically based on the script; edit it if it isn't right.

Click Add to Media Library to keep it. Until you do, the audio is only a preview.

Trying several versions

Generating again doesn't throw away what you had. The History tab keeps the previews you've made, so you can flip between two readings of the same line and pick the better one. Selecting one from History puts its text, voice and name back into the form, which is the quickest way to make a small change to something you nearly liked.

Clear form empties everything and starts fresh.

Things worth knowing

  • Generating audio uses your organisation's AI credits.

  • Punctuation does a lot of work. Commas and full stops are how you control pacing — a script written as one long sentence will be read as one long sentence.

  • Once the audio is in your media library it behaves like any other audio file: you can add it to an audio activity, add subtitles to it, and manage it from the library.

Did this answer your question?