SKYHOUSE.dev Journal

Maintaining the Cloud Fortress

inspiredby: Haiku reader timer, people without Wikipedia

Why: Damien wanted the long tail of song subjects (grandmothers, exes, friends) and found the regex extractor's claims spotty; when we measured it, it was right 8% of the time on pages that mention an inspiration, while Claude Haiku was right 92%.

A day of research into better data collection ended in three server-side changes: a new hourly systemd user timer that has Claude Haiku read Wikipedia song articles, a schema migration so the site can list people and songs that have no Wikidata item, and a fix so artists get photos. Everything was deployed from ~/dev/inspiredby with scheduler/deploy.sh after a backup.

1. The LLM reader timer

I added inspiredby-reader.service + .timer (user units, hourly, not persistent). Each run calls run.py read, which takes up to 100 queued song/album articles from the local Wikipedia dump and sends each to Haiku through claude -p (Claude Code headless, under Damien's own login, no tools, page text fenced as untrusted). Answers are only kept when their quote is really in the article and names the person. A page the reader has read replaces the regex extractor's claims for it.

2. Schema migration: people without Wikipedia

The production DB went to user_version 16: people.local_key and the editor table person_alias. People with no Wikidata item are keyed per artist (local:<artist>:<name>) and must be named. Backup taken first: ~/inspiredby/backups/inspiredby-2026-10-06-pre-step0.db. I also imported 68 web-research leads (from Sonnet research agents) for the worker to verify.

3. Artist photos

The photo job only ever queued people somebody wrote about, so 3,881 artists (Eels, Florence + the Machine…) had none. Now it queues them too; the worker downloaded about 4,000 more thumbnails into ~/inspiredby/media/photos/ the same evening.

Net effect: one more hourly user timer, using Claude rather than Wikimedia; no new ports, no cron changes.

← Back to Admin Hub