Text-to-Speech and Speech-to-Text
Text-to-Speech (TTS) and Speech-to-Text (STT) split into two paths according to Type.
- On Device: Converts with a neural-network model built into the controller, requiring neither an internet connection nor an API key; it responds quickly and audio never leaves the device. Selecting a Language lists the Model options for it, and the chosen model downloads automatically on first use.
- Google Cloud, OpenAI, ElevenLabs: Converts through an online service and needs an API Key, which is stored encrypted on the controller rather than on the Grablo server. TTS supports all four types; STT supports On Device, Google Cloud, and OpenAI. Google Cloud offers many languages and voices, ElevenLabs offers emotionally expressive voices, and OpenAI STT uses the Whisper model. See the Appendix for how to obtain each API key.
The On Device Language list has ten entries for TTS (English, Korean, Chinese, German, French, Spanish, Portuguese, Hindi, Arabic, Russian) and eight for STT (English, Korean, Chinese, Japanese, German, Spanish, Portuguese, Russian); English is the default for both. Models are selected by the original name shown on screen, and the list and default for each language are as follows.
On Device TTS models
| Language | Models (default in bold) |
|---|---|
| English | vits-piper-en_US-lessac-low, vits-piper-en_US-lessac-medium, vits-piper-en_US-lessac-high, vits-piper-en_US-ryan-low, vits-piper-en_US-ryan-medium, vits-piper-en_US-ryan-high, kitten-nano-en-v0_2-fp16, vits-melo-tts-zh_en |
| Korean | vits-mimic3-ko_KO-kss_low |
| Chinese | sherpa-onnx-vits-zh-ll, vits-piper-zh_CN-huayan-medium, vits-melo-tts-zh_en |
| German | vits-piper-de_DE-thorsten-low, vits-piper-de_DE-thorsten-medium, vits-piper-de_DE-thorsten-high, vits-piper-de_DE-kerstin-low |
| French | vits-piper-fr_FR-siwis-low, vits-piper-fr_FR-siwis-medium, vits-piper-fr_FR-tom-medium, vits-piper-fr_FR-gilles-low |
| Spanish | vits-piper-es_ES-davefx-medium, vits-piper-es_MX-claude-high |
| Portuguese | vits-piper-pt_BR-faber-medium, vits-piper-pt_BR-edresson-low |
| Hindi | vits-piper-hi_IN-pratham-medium, vits-piper-hi_IN-priyamvada-medium |
| Arabic | vits-piper-ar_JO-kareem-low, vits-piper-ar_JO-kareem-medium |
| Russian | vits-piper-ru_RU-denis-medium, vits-piper-ru_RU-irina-medium |
On Device STT models
| Language | Models (default in bold) |
|---|---|
| English | sherpa-onnx-whisper-tiny, sherpa-onnx-whisper-base, sherpa-onnx-whisper-small, sherpa-onnx-whisper-medium, sherpa-onnx-nemo-ctc-en-conformer-small, sherpa-onnx-nemo-ctc-en-citrinet-512, sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17 |
| Korean | sherpa-onnx-zipformer-korean-2024-06-24, sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17 |
| Chinese | sherpa-onnx-paraformer-zh-int8-2025-10-07, sherpa-onnx-paraformer-zh-small-2024-03-09, sherpa-onnx-zipformer-ctc-small-zh-int8-2025-07-16, sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17 |
| Japanese | sherpa-onnx-zipformer-ja-en-reazonspeech-2025-01-17, sherpa-onnx-zipformer-ja-reazonspeech-2024-08-01, sherpa-onnx-sense-voice-zh-en-ja-ko-yue-2024-07-17 |
| German | sherpa-onnx-whisper-tiny, sherpa-onnx-whisper-base, sherpa-onnx-whisper-small, sherpa-onnx-whisper-medium, sherpa-onnx-nemo-stt_de_fastconformer_hybrid_large_pc-int8 |
| Spanish | sherpa-onnx-whisper-tiny, sherpa-onnx-whisper-base, sherpa-onnx-whisper-small, sherpa-onnx-whisper-medium, sherpa-onnx-nemo-fast-conformer-ctc-es-1424-int8 |
| Portuguese | sherpa-onnx-whisper-tiny, sherpa-onnx-whisper-base, sherpa-onnx-whisper-small, sherpa-onnx-whisper-medium, sherpa-onnx-nemo-stt_pt_fastconformer_hybrid_large_pc-int8 |
| Russian | sherpa-onnx-whisper-tiny, sherpa-onnx-whisper-base, sherpa-onnx-whisper-small, sherpa-onnx-whisper-medium, sherpa-onnx-zipformer-ru-int8-2025-04-20, sherpa-onnx-small-zipformer-ru-2024-09-18 |
Selecting CUSTOM in the model list lets you enter the path to a model folder you prepared yourself in Model folder. Renaming the downloaded folder prevents the model from being recognized, so keep its original name. Models are downloaded from the sherpa-onnx release pages (TTS, STT), and each model is described in the sherpa-onnx documentation (TTS, STT). For STT, choosing CUSTOM also shows a Language text field (default en); the language format varies by model, so follow that model's documentation.
The Google Cloud type selects a Language from Korean, English (US) (the default), English (UK), Japanese, Chinese, German, French, Italian, and Spanish. OpenAI TTS selects a Model: tts-1 (the default, optimized for speed and suited to real-time use) or tts-1-hd (optimized for audio quality, more natural and expressive but comparatively slower).
The Speech-to-Text Input Source is Microphone (Input) (the default) or Speaker (Output). Choosing the speaker transcribes sound played by the device itself, such as media playback or voice prompts.
Speech-to-Text has two settings that decide where speech starts and ends. Voice Detection Threshold ranges from 5 to 90% (30% by default): lowering it picks up quieter sounds but also background noise (20 to 40% for quiet rooms, 30 to 50% for normal ones, 50 to 70% for noisy ones). Speech Completion Timeout is the silence after which speech is treated as finished, from 0.3 to 5 seconds (1 second by default; 0.5s for fast exchanges, 1s for normal speech, 2s for deliberate speech).