Generate and enhance speech in Creative Studio — text-to-speech with ElevenLabs or Gemini voices, dialogue generation, voice cloning, and audio enhancement.
Audio nodes produce speech and audio assets for ads, Hyperframes voiceovers, URL-to-video handoffs, and standalone Studio video pipelines. Add an audio node to any canvas, choose a mode, and connect the output to downstream video or export nodes.
Convert a script to natural-sounding speech. Paste or type your text, choose a voice, and generate.Default engine: ElevenLabs V3. Gemini TTS is also available as an alternative.Best for: voiceovers, ad narration, product explainers, Hyperframes scripts.
Generate multi-turn or conversational speech — two or more distinct speakers in sequence, useful for interview-style or call-and-response content.Best for: podcast-style ads, scripted conversations, multi-character scenes.
Browse and preview available voices before committing to a full generation run. Sample any voice with a short test phrase.Best for: finding the right voice tone and accent for your brand or campaign.
Upload a reference audio sample to create a custom voice that sounds like a specific person or persona. Credits apply for clone creation.Best for: consistent brand voice, spokesperson content, localized campaigns with a familiar voice.
Clean up existing audio — reduce noise, normalize levels, and improve clarity. Feed any uploaded audio file or upstream audio node output.Best for: improving recorded VO, cleaning location audio, prepping assets for lip sync.
Open the toolbar and add an Audio node. Choose TTS or Dialogue mode.
2
Write your script
Type or paste the script into the text field. For Dialogue mode, label each speaker turn.
3
Select a voice
Use the Voice selector mode to preview options, then switch back to TTS and apply your chosen voice.
4
Generate
Click Run. The audio renders as a waveform output on the node.
5
Connect to video or export
Drag the audio output to a lip sync node, a Hyperframes voiceover slot, or an export/merge node.
For Hyperframes, generate voiceover at the Voiceover stage rather than attaching a separate audio node — Hyperframes handles timing alignment automatically when you use the built-in voiceover step.