Login
Free Sign Up
Docs
/

AI - Text to Speech

Type identifier: ai:text-to-speech Category: AI Operations

Description

The AI - Text to Speech node generates realistic spoken audio from text input using OpenAI's text-to-speech models. It supports multiple voice options and quality levels, returning an audio file that can be used in Experiences or stored.

Input Handles

Handle

Type

Description

text

string

The text to convert to speech.

model

string

(Dynamic) The TTS model to use.

voice

string

(Dynamic) The voice to use.

Note: model and voice handles appear when set to dynamic mode.

Output Handles

Handle

Type

Description

audio

file

The generated audio file.

Configuration Options

Text Input

Option

Type

Default

Description

Text

string or dynamic

dynamic

The text content to synthesize.

Model Selection

Option

Type

Default

Description

Model

literal or dynamic

openai:tts-1

The TTS model to use.

See AI - Model Reference — Text to Speech Models for the canonical model list.

Available models:

Model

Quality

Speed

Use Case

openai:tts-1

Standard

Fast

Real-time applications

openai:tts-1-hd

High Definition

Slower

High-quality output

Voice Selection

Option

Type

Default

Description

Voice

literal or dynamic

openai:alloy

The voice character to use.

Available voices:

Voice

Character

openai:alloy

Neutral, balanced

openai:echo

Warm, natural

openai:fable

Expressive, British

openai:onyx

Deep, authoritative

openai:nova

Friendly, energetic

openai:shimmer

Clear, pleasant

Behaviour

  1. Resolves text content from input
  2. Generates audio with the selected model and voice
  3. Stores the audio as a file in the Space
  4. Returns the audio file

Audio Format

Generated audio is typically in MP3 format.

Text Limits

OpenAI TTS has a character limit per request. Long text may need to be chunked.

Examples

Basic Text to Speech

Configuration:

  • Model: openai:tts-1
  • Voice: openai:nova

Input (text):

"Welcome to our application. Let me guide you through the features."

Output: MP3 audio file with spoken content.

High-Quality Narration

Configuration:

  • Model: openai:tts-1-hd
  • Voice: openai:onyx

Use case: Generate professional-quality audio for podcasts or presentations.

Dynamic Voice Selection

Configuration:

  • Model: literal openai:tts-1
  • Voice: dynamic

Use case: Allow user to select their preferred voice at runtime.

Related pages