Login
Free Sign Up
Docs
/

AI - Transcribe

Type identifier: ai:transcribe Category: AI Operations

Description

The AI - Transcribe node converts spoken audio into text using OpenAI's Whisper model. It supports various audio formats and can handle multiple languages.

Input Handles

Handle

Type

Description

audio

file

The audio file to transcribe.

Output Handles

Handle

Type

Description

transcription

string

The transcribed text content.

Configuration Options

Option

Type

Default

Description

Model

"openai:whisper-1"

"openai:whisper-1"

The transcription model to use.

See AI - Model Reference — Transcribe Models for the canonical model list.

Behaviour

  1. Receives audio file from input handle
  2. Transcribes the audio with the selected model
  3. Returns the transcribed text

Supported Audio Formats

  • MP3
  • MP4
  • MPEG
  • MPGA
  • M4A
  • WAV
  • WEBM

Language Support

Whisper automatically detects the spoken language and provides transcription. It supports multiple languages with varying levels of accuracy.

File Size Limits

Audio files have size limits imposed by the API. Large recordings may need to be split.

Examples

Basic Audio Transcription

Input: Audio recording of a meeting

Output (transcription):

"Good morning everyone. Today we'll discuss the Q3 roadmap.
First, let's review the progress from last quarter..."

Voice Message Processing

Flow:

  1. UX - Action receives audio from user
  2. AI - Transcribe converts to text
  3. AI - Chat Message processes the transcription

Multi-Language Support

Use case: Transcribe customer support calls in various languages, then translate or analyse the content.

Related pages