Speak-Summary
Speak-Summary is an code AI skill with a core value of Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. It
helps developers solve real-world problems in the code domain, boosting
efficiency, automating repetitive tasks, and optimizing workflows.
Convert text, markdown, or a summary produced by another skill into a listenable MP3 using local CPU-only neural text-to-speech. Rewrites written prose for the ear before synthesising. Use when the us
Quick Facts
mkdir -p ./skills/speak-summary && curl -sfL https://raw.githubusercontent.com/github/awesome-copilot/main/skills/speak-summary/SKILL.md -o ./skills/speak-summary/SKILL.md Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).
Skill Content
# Speak Summary
Turn written text into audio someone will actually want to listen to.
This skill is deliberately a **terminal step in a chain**. Another skill (or you)
produces the text; this one makes it listenable. It pairs naturally with
`roundup`, `daily-prep`, `meeting-minutes`, or any summarisation work.
Everything runs locally on CPU. No text is sent to a cloud speech service, which
matters when the content is confidential, and it means the skill works in a
headless cloud agent or CI container just as well as on a laptop.
Prerequisites
The synthesis engine is [Kyutai `pocket-tts`](https://github.com/kyutai-labs/pocket-tts),
a small neural TTS model designed to run on CPUs.
The bundled script installs it automatically into a cached virtualenv on first
use, so usually you need do nothing. To install it explicitly:
pip install pocket-tts # any platform
brew install pocket-tts # macOS, if preferred`pocket-tts` requires **Python >=3.10 and <3.15**. The script searches for a
compatible interpreter rather than assuming `python3` is one — worth knowing if
you are on a very new Python, where installation would otherwise fail.
You also need an encoder. `ffmpeg` is strongly preferred (`brew install ffmpeg`
or `apt-get install -y ffmpeg`); on macOS the script falls back to the built-in
`afconvert` and emits `.m4a` instead of `.mp3`.
The first run downloads the model (~1GB) from Hugging Face. After that it is
fully offline and synthesises roughly 6x faster than real-time.
The important step: rewrite for the ear
**Do not feed written text straight into the synthesiser.** Prose that reads well
on screen is tiring to listen to. Rewriting it first is what separates a useful
audio digest from an unlistenable one.
Produce a spoken script that:
- **Opens with orientation.** What this is, what it covers, roughly how long it runs.
- **Replaces bullets with connective prose.** "First… The bigger one is… Finally…" — a listener has no visual structure to lean on, so carry it in the language.
- **Expands abbreviations on first use.** "PR" becomes "pull request", "CI" becomes "continuous integration". Acronyms that read fine are noise when spoken.
- **Speaks dates and numbers naturally.** "the twentieth of August", not "2026-08-20". "About three thousand", not "2,847".
- **Never reads URLs aloud.** Say "linked in the written version" instead.
- **Uses short sentences.** Split anything past roughly 25 words.
- **Signposts transitions.** "Turning to the product side…", "Two things need your attention…".
- **Ends with the actions.** Recap what the listener should do, since that is what they need to retain and they cannot scroll back.
- **Drops anything purely visual.** Tables, code blocks, and diagrams should be summarised in a sentence or omitted, never read out.
Write this spoken script to its own `.txt` file. Keep the original written
version with its links intact — the audio is a companion to it, not a
replacement. The user will want to click through later.
Synthesise
./scripts/tts.sh <input.txt> <output.mp3> [voice.safetensors]The script strips any residual markdown, splits the text on sentence boundaries
into ~600 character chunks (quality degrades on long single inputs), synthesises
each chunk, and concatenates the result into a mono MP3 at 96kbps — small enough
to sync to a phone, good enough for speech.
Environment overrides:
| Variable | Purpose |
|---|---|
| `SPEAK_TTS_BIN` | Path to a specific `pocket-tts` binary; skips all auto-detection. |
| `SPEAK_TTS_HOME` | Where to create/find the cached virtualenv. Default `~/.cache/speak-summary/venv`. |
Voices
The default English voice is `alba`. To use a different one, `pocket-tts`
supports voice cloning from a short clean audio sample:
pocket-tts export-voice --helpPass the resulting `.safetensors` file as the third argument to the script.
Only clone a voice you have the rights to use. Do not clone a r
🎯 Best For
- GitHub Copilot users
- Claude users
- Software engineers
- Development teams
- Tech leads
💡 Use Cases
- Code quality improvement
- Best practice enforcement
📖 How to Use This Skill
- 1
Install the Skill
Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.
- 2
Load into Your AI Assistant
Open GitHub Copilot or Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.
- 3
Apply Speak-Summary to Your Work
Open your project in the AI assistant and ask it to apply the skill. Start with a small module to verify the output quality.
- 4
Review and Refine
Review AI suggestions before committing. Run tests, check for regressions, and iterate on the skill output.
❓ Frequently Asked Questions
Is Speak-Summary compatible with Cursor and VS Code?
Yes — this skill works with any AI coding assistant including Cursor, VS Code with Copilot, and JetBrains IDEs.
Do I need specific dependencies for Speak-Summary?
Check the install command and Works With section. Most code skills only require the AI assistant and your codebase.
How do I install Speak-Summary?
Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/speak-summary/SKILL.md, ready to use.
Can I customize this skill for my team?
Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.
⚠️ Common Mistakes to Avoid
Skipping validation
Always test AI-generated code changes, even for simple refactors.
Missing dependency updates
Check if the skill requires updated dependencies or new packages.