Skip to main content
ElevenLabs provides speech-to-text transcription and high-quality text-to-speech capabilities for your agent.

Voice transcription

All audio transcription on the platform uses ElevenLabs Scribe — no setup or API key needed. How to use: Press the mic button in the chat to dictate, or upload an audio or video file (for example a meeting recording) to an agent.
  • Mic button: your speech becomes clean text — filler words are stripped, and you can set your preferred transcription language in your user settings.
  • Uploaded recordings: the transcript includes who said what and when, with each speaker labelled and timestamped:
Example use cases:
  • Talk instead of typing to your agent
  • Drop a meeting recording and ask the agent to summarize decisions per person
See Voice Communication for details on voice input and walk-and-talk mode.
Workspace admins can decline ElevenLabs under Workspace settings → Data processors. Transcription then falls back to OpenAI Whisper — plain text without speaker labels or timestamps.

Generated voiceovers

Go to Settings → Capabilities and enable Text to Speech. ElevenLabs is one of the available voice providers — your agent can pick a specific ElevenLabs voice per request when generating audio files as a tool (for example, an audio summary embedded in a newsletter).
OpenAI’s text-to-speech is also available as an alternative. Both options are included in the Text to Speech capability.
Example use cases:
  • “In the weekly newsletter, include an audio clip of the executive summary”
  • “Generate a voiceover for this script using a male Scottish voice”
  • “Every morning, generate a ‘daily Mr President’ style update that I can listen to while driving to work”

Agent voice

ElevenLabs voices also show up as options for the agent’s own voice — the one used by the Speak button and Walk & Talk mode. Premade, professional, and cloned voices are all available, alongside OpenAI voices, in a single dropdown under Settings → Basic Information → Voice.