Speech generation
Turn text into WAV audio files in the agent folders, through OpenRouter (Gemini 3.8 Flash TTS or Flash Lite TTS). Uses the workspace OpenRouter key.
Turn text into WAV audio files in the agent folders, through OpenRouter (Gemini 3.8 Flash TTS or Flash Lite TTS). Uses the workspace OpenRouter key.
Slug: speech-generation
Unlocks tools: generate_speech
The rest of this page is the exact guidance this skill injects into an agent's system prompt.
Speech generation
This tool group gives you generate_speech: it turns text into a WAV audio file written in one of your folders, and returns its path. (WAV, not mp3: the models answer raw audio.)
How to call it
text: the exact words to speak, at most 5,000 characters. Everything in it is read aloud: no stage directions, no "(pause)", no "[cheerful]".style: the tone, in a few words ("warm and friendly", "calm and slow", "whispering"). It is not read aloud.path: where to write the file, ending in.wav(e.g.audio/intro.wav). Missing folders are created.model:google/gemini-3.8-flash-tts(default, the expressive tier) orgoogle/gemini-3.8-flash-lite-tts(faster and cheaper, plain reading).voice: the voice name as Google names it;Koreis the default.
Discipline
- Longer than 5,000 characters: split at paragraph boundaries into several files (
part-1.wav,part-2.wav…) and say so in your answer. - Give the path back in your answer: the owner finds the file there.
- A failure is said, not retried blindly. No OpenRouter key, an unknown voice, a provider error: report the reason the tool returned.
Command execution
Run shell commands in the agent workspace (install deps, run scripts, build steps, CLIs). Every command requires your approval by default; an optional per-agent "Yolo" mode auto-runs them.
Coding CLI (Claude Code / Codex)
Delegate complete dev tasks (analyse, review, debug, implement) to the coding CLI installed on this machine — Claude Code or Codex — running under the owner's subscription. Read-only by default; every run requires approval unless Yolo is enabled.