Get Started
Text-to-Speech (TTS)

Text-to-Speech (TTS)

Converts text to speech and plays it. The service type (On Device, Google Cloud, OpenAI, or ElevenLabs) is set when the service is registered in Settings > Text-to-Speech (Part 6), and the field that selects the voice depends on the type. See Appendix B for obtaining API keys for the online services and the ElevenLabs Voice ID.

Text-to-Speech

Select the text-to-speech service to use.

Commands

  • Play (default): Converts the text to speech and plays it.
  • Stop: Stops the speech being played.

Fields

Speaker ID, Voice ID

When the type is On Device, enter the model's speaker ID as a number. IDs are defined per model and start from 0; an invalid value selects the first voice. The speaker IDs of each model are listed in the sherpa-onnx model list. When the type is ElevenLabs, enter the Voice ID of the voice to use; this field is displayed only when Command is Play.

Voice Name

When the type is Google Cloud, enter the voice name. The default is en-US-Standard-C, and the voices are listed in the Google Cloud documentation. When the type is OpenAI, select alloy (default), echo, fable, onyx, nova, or shimmer; the voices are described in the OpenAI documentation. In both cases the field is displayed only when Command is Play.

Output Text

Enter the text to convert to speech. Displayed only when Command is Play.

Wait for Completion

Select whether the action stays running until the speech ends. The default is on. Displayed only when Command is Play.

Playback Speed

Select 0.25x, 0.5x, 0.75x, 1.0x (default), 1.25x, 1.5x, 2.0x, 3.0x, or 4.0x, or select Manual Input and enter 0.1 to 10 in Playback Rate (x). Displayed when Command is Play, and not displayed when the type is ElevenLabs.