Glanceable podcast
A Claude Code skill that turns YouTube videos and web articles into glanceable podcast episodes in your own private podcast feed. Articles are read aloud with open-source text-to-speech and YouTube videos are converted to audio MP3 files with frequently updating chapter artwork from the video itself.
Overview
Some things only exist as videos but work fine as audio: talks, lectures, interviews. Alex Chan’s glancecast (original tool) converts a video into an MP3 and embeds a frame every few seconds as an ID3 chapter image. Podcast apps that show chapter artwork (Overcast, Pocket Casts) then play it as a slow slideshow you can glance at while you listen.
The same idea works for articles: the text is read aloud, and each of the article’s images becomes chapter artwork at the point the text reaches it.
This tool is packaged as a Claude Code skill. Give Claude a link and:
- For a video, it downloads it with
yt-dlp, extracts the audio and frames withffmpeg(every n seconds, or one per scene change for slides), writes the frames into the MP3 as chapter images, and writes show notes and chapter timestamps from the transcript. - For an article, it extracts the text and images, fixes how names and awkward captions will sound, then reads it aloud with Kokoro (Apache-2.0). Captions are read in a second voice after a soft chime, and headings become chapter timestamps.
- Either way, it uploads the episode to a Cloudflare R2 bucket and adds it to an RSS feed at a secret URL.
Setup
Follow the six steps on the tool page: install ffmpeg, uv and deno; install the skill; create an R2 bucket and a bucket-scoped API token; store its keys in macOS Keychain or 1Password; create the feed; subscribe.
Usage
In Claude Code:
Make this a glanceable podcast episode: https://www.youtube.com/watch?v=…
Add this talk to my glanceable podcast feed, one picture per slide.
Read this later as a podcast: https://example.substack.com/p/…
What’s in my glanceable podcast feed? Remove the one about …
What’s in the skill
| File | Purpose |
|---|---|
skill/SKILL.md |
Instructions for Claude: both workflows, frame-mode choice, article review rules, show notes, secret handling, troubleshooting |
skill/scripts/video_episode.py |
Video → episode.mp3 with chapter images, cover.jpg, episode.json and an optional transcript.txt |
skill/scripts/fetch_article.py |
Article URL (or saved .html) → article.json, images fitted to 945px, grids for big galleries, cover |
skill/scripts/article_episode.py |
article.json → spoken episode.mp3 with image chapters, and episode.json |
skill/scripts/publish.py |
init, check, add, list, remove, feed-url and rotate for the feed on R2 |
skill/scripts/glance.py |
Helpers shared by the episode scripts (chapter writing, sizes) |
The scripts are self-contained uv scripts with inline dependencies (yt-dlp, mutagen, boto3, beautifulsoup4, readability-lxml, pillow, kokoro), so there’s nothing to pip install. The work is done by the scripts; Claude only handles the parts that need judgement (frame mode, titles, notes, sections, pronunciations, spoken captions), so each run is consistent.
Episode files
- Audio: mono 64 kbps MP3 by default (
--stereo,--bitrateto change). - Frames: scaled to fit 945×945 (the largest chapter art the iPhone shows). Each frame is a
CHAPframe with anAPICimage. There’s deliberately noCTOC(table of contents), so players don’t list hundreds of 5-second chapters. - Scene mode (
--scene 0.3): one frame per scene change. A burst of changes during a transition becomes a single chapter showing the settled frame.
Article episodes
- Extraction: a dedicated extractor for Substack (full-size images, captions, galleries; subscribe widgets and footnote markers dropped), and Readability for other sites. Paywalled or JavaScript-only pages can be passed as a saved
.htmlfile. - Voices: Kokoro-82M, chosen in a bake-off against Chatterbox and Kyutai TTS for sounding most natural and being by far the fastest (about 10× real time on an M1 Pro). The article is read by
af_heart; captions byam_puckat 1.1×. - Images: each image becomes a chapter at the point the text reaches it, with a soft two-note chime (generated, not a sample) and its caption. Back-to-back images get one chime, then a quieter tick for each, and are held on screen for at least 4 seconds each. Galleries of more than four images become one grid. Images with no caption or useful alt text get the chime alone.
- Before the first image, the article’s share image (or a generated title card) is shown while the title and byline are read.
- Maths: formulas (MathML, as MathJax and KaTeX write it) are rendered to images with MathJax in Chrome. Display equations become silent picture chapters; inline formulas are named when short ("Delta M", "q prime") and said as "shown" otherwise, while a card of the paragraph's formulas is on screen.
- Journals and saved pages: pages on the Atypon platform (PNAS and others) get their own extractor, with the abstract and theorem boxes. Sites that block scripts can be saved from the browser and passed as a file.
- Pronunciation: Claude adds respellings (
"McKnight": "Mick-Nite") or Kokoro phonemes for names it would get wrong, and rewrites captions so they sound natural aloud, without changing what they say.
Feed layout on R2
feeds/<feed-token>.xml RSS feed (itunes:block, so directories don't list it)
feeds/<feed-token>.json episode list the feed is rebuilt from
episodes/<random>.mp3 audio
episodes/<random>.jpg episode artwork
covers/<random>.jpg feed artwork
Every path is random, so knowing one episode’s URL doesn’t reveal the feed or the other episodes.
Secrets
There are three secrets: the two R2 access keys, which you add, and the feed token, which publish.py init generates and saves alongside them. They live in a secret store on your Mac, chosen with publish.py init --secrets:
| Store | Adding the R2 keys | init options |
|---|---|---|
| macOS Keychain (default) | security add-generic-password -s glanceable-podcast -a <name> -w, which prompts for the value so it never reaches shell history |
none |
| 1Password | In the 1Password app, add password fields labelled r2_access_key_id and r2_secret_access_key to an item (default title “Glanceable podcast”). The scripts read them with the op CLI through its desktop-app integration, so access is approved with Touch ID. |
--secrets 1password --op-vault <vault> [--op-item <title>] [--op-account <account>] |
With 1Password, use the vault's name as the CLI shows it (op vault list). In Teams and Business accounts the built-in vault the app calls “Employee” is named Personal in the CLI. If the CLI is signed into more than one account (op account list), add --op-account, e.g. --op-account example.1password.com.
Whichever store you use:
~/.config/glanceable-podcast/config.jsonholds only non-secret settings: account ID, bucket, public URL, feed title and which secret store to use.- The skill tells Claude never to read or print secrets or the feed URL.
publish.py feed-urlcopies the URL to the clipboard instead of printing it. - The R2 token should be scoped to Object Read & Write on this one bucket. If the feed URL leaks,
publish.py rotatemoves the feed to a new token.
Nothing secret lives in this repo. The site publishes only the instructions and scripts.
Credit
The video technique is based on Alex Chan’s glancecast project. Alex had the idea of making videos into glanceable podcasts. Thanks, Alex!