NewSnippets: tell the AI about you and your company once.See how New modelClaude Sonnet 5.5 is now available.Read more
CREATE

The voice-over and the soundtrack, in one place.

Store announcements, training audio, a narration for the product video, a jingle for the launch. Type the script or describe the track, pick the model, and press play. Both apps sit in the same workspace as your chats, with the same models, usage record and admin controls.

Text to Speech Done
Northwind launch update
Eleven v3 · Darian, warm grounded storyteller
0:00 / 0:28

Text Hi team, quick update before Tuesday's launch. Northwind Analytics is feature-complete and the Seattle sync issue is fixed...

Music Generation Done
Product launch background track
Lyria 3 Clip · MP3
0:00 / 0:30

Prompt Upbeat, optimistic product-launch background track: warm electric piano, tight modern drums, plucky synth bass, subtle handclaps, 112 BPM, confident and clean, no vocals.

Both generated in StickyPrompts from the text and prompt shown. Unedited.
Two audio apps

Speech and music, without a studio booking.

Text to Speech turns a script into a finished read. Music Generation turns a sentence into a track. Both give you a file you can use the same afternoon.

Voices with range
Eleven v3 is the default, built for emotionally rich delivery. Gemini TTS models sit next to it when you want a different read of the same script.
Voice and language, your call
Pick a voice from the list and set the language, or leave it on auto-detect and let the text decide. The same script can go out in every language your stores speak.
Play, copy, download
Every result opens with a player, the exact text it was made from and a Download button. No export step, no separate audio tool.
Music from a sentence
Describe the mood, the instruments and the tempo. Eleven Music v2 is the default, Lyria 3 Pro writes full songs, Lyria 3 Clip makes short clips and loops.
Length and vocals you set
Choose the duration anywhere from 3 to 300 seconds, pick the output format, and switch Instrumental on when the track sits under a voice.
A record of every take
Each generation keeps its model, voice or prompt, cost and date, so the version the team approved is easy to find again.
TEXT TO SPEECH

Paste the script, pick the voice.

Type or paste the text, choose the model, the voice and the language, and generate. Eleven v3 is selected by default for its expressive delivery; switch to a Gemini TTS model when you want a different sound or a quick draft read.

  • Eleven v3, Gemini 3.8 Flash TTS, Flash-Lite TTS and more
  • A voice list with a short description of each voice
  • Set the language, or leave it on auto-detect
Text to speech: choose the model, voice and language 1 Eleven v3 by default 2 Choose the voice 3 Or set the language
THE RESULT

Listen, check the words, download.

The result opens with a player, the exact text that was read and the details of the take: model, voice, language, cost and date. If the read is right, download it. If it is not, change a line and generate again.

  • Player and Download on the same screen
  • The source text kept next to the audio
  • Model, voice and cost recorded with each take
Generated speech ready to play and download 1 Play it back 2 Download the file 3 Voice and cost, recorded
A week of audio at Northwind

Five jobs that used to need an agency.

Northwind runs retail stores and a support line. This is what its teams make with the two apps in an ordinary week.

Operations

Store announcements

Opening hours, a click-and-collect reminder, the weekend offer. Written once, voiced once, played in every store the same way.

Text to Speech · Eleven v3
Marketing

Product video voice-overs

A 30-second narration for the new autumn range, ready to lay under the cut. Try a second voice before anyone books a studio.

Text to Speech · Gemini TTS
Store training

Training audio

The till procedure and the returns policy as short audio lessons new starters can play on shift, in their own language.

Text to Speech · language set per script
Marketing

A launch jingle

A short, bright sting with the brand's energy for the Northwind Analytics launch video and the social cut-downs.

Music Generation · Lyria 3 Clip
Customer Support

On-hold music

A calm instrumental loop for the support line, long enough that nobody hears it restart while they wait.

Music Generation · Eleven Music v2, instrumental
MUSIC GENERATION

Describe the track, set the length.

Write what you hear in your head: the mood, the instruments, the tempo, where it will be used. Pick the model, the output format and the duration, turn Instrumental on if it sits under a voice, and generate.

  • Eleven Music v2 by default, plus Lyria 3 Pro and Lyria 3 Clip
  • Duration from 3 to 300 seconds
  • Instrumental toggle for beds and hold music
Music generation: Eleven Music and Lyria models, duration and instrumental options 1 Full songs with Lyria 3 Pro 2 3 to 300 seconds 3 Vocals on or off
THE TRACK

The track, the prompt and the lyrics.

Every track comes back with a player, the lyrics if it has any, the prompt it was made from and a Download button. The prompt is kept, so when marketing asks for the same feel at 60 seconds, you start from what worked.

  • Play in the browser, download the MP3
  • Lyrics and prompt, each one click to copy
  • Model, duration, format and cost on every track
Pair it with video generation
A generated product-launch track 1 Play the track 2 The prompt, kept 3 Model, length and cost
WHY AUDIO BELONGS IN THE WORKSPACE
  • Write the script in chat, voice it in Text to Speech, without switching accounts
  • Usage metered like every other model, on the same plan
  • One login and one set of controls for text, images, video, speech and music
  • Every take keeps its text or prompt, so the approved version is easy to find
  • Voice-overs and music for the video you made in the same place
  • No separate audio tool to buy, onboard or offboard
Straight answers

What teams ask about voice and music.

Which models can we use?

For speech, Eleven v3 is the default, with Gemini TTS models alongside it. For music, Eleven Music v2 is the default, with Lyria 3 Pro for full songs with verses, choruses and vocals, and Lyria 3 Clip for short clips and loops.

How long can a track be?

You set the duration on a slider from 3 to 300 seconds before you generate. Lyria 3 Clip is made for short pieces of around 30 seconds; pick Eleven Music v2 or Lyria 3 Pro for anything longer.

Can I get music without vocals?

Yes. Turn on Instrumental and the track comes back without singing, which is what you want under a voice-over or on a phone line. Where a track has lyrics, they are shown next to the player so you can copy them.

What do I get at the end?

A result page with a player, the text or prompt it was made from, the model and settings, the cost and a Download button. Nothing to export, and the take stays in the workspace for the next person who needs it.

How is it billed?

Like everything else in the workspace: each generation is metered as AI usage at the provider's rate plus your plan's commission, and the cost is shown on the result. There is no separate audio subscription to buy or manage.

Trial
A $5 balance to start. No card needed.
Start with one announcement

Type this week's store message and hear it read.

Start free with a $5 trial balance. Try a few voices on the same script and keep the one that sounds like you.