Local Wikipedia copy, and inspiredby gathering again
Why: inspiredby's task queue had run dry, so the hourly worker was waking up, finding nothing, and exiting. Crawling Wikipedia's ~250k song and album articles at one request every 2 seconds would have taken weeks, so we downloaded the whole thing instead.
Two related changes. First, a full English Wikipedia text dump now lives at /home/plex/data/wikipedia/, with a small stdlib-only library and an SQLite index, meant to be shared by any project on the box (FILM FUTURES is next). Second, inspiredby (song → real-person inspiration index, inspiredby.skyhouse.dev) now reads from that copy, plus a few new sources, and its worker's hourly budget changed.
1. Local Wikipedia at ~/data/wikipedia
I downloaded the 2026-10-01 enwiki pages-articles-multistream dump (26.9 GB compressed, text only, no media), its stream index, and the page_props table (page → Wikidata Q-id). All four files were verified against Wikimedia's published SHA-1s. wikidump.py build then indexed it in ~17 minutes on 11 cores (peak RSS ~1.3 GB): 7.25 M articles (99.9% with Q-ids), 2.65 M category pages, and every page's infoboxes and categories. Pages are read on demand by seeking into the bz2 file, ~5–60 ms per page.
~/data/wikipedia/wikidump.pyis the library + CLI.README.mdthere is the how-to. The folder is its own little git repo (dumps and.dbignored).current -> 20261001;20261001/wikidump.dbis the index (8.0 GB). Total footprint ~34 GB on the root SSD (675 GB free afterwards).- Readers must check
meta.complete = 1. A rebuild writeswikidump.db.buildingand renames it only at the end, so readers never see half an index. - Gotcha found on the first full build: 21 titles appear twice, because pages were renamed while Wikimedia was writing the dump. The build now keeps the real article over the leftover redirect.
- Handoff prompt for other projects:
~/data/wikipedia/PROMPT-film-data.md(SSH asplexon port 2222, run Python on the server, treat the folder as read-only).
2. inspiredby: new sources and a different worker budget
The worker unit (inspiredby-worker.service, user unit, hourly timer) now runs run.py work -n 250, where 250 counts only tasks that touched the network. Tasks served from the HTTP cache or the local dump are free, and every run stops after 50 minutes so it can't overlap the next. In practice it now uses the CPU for most of each hour while it works through ~246k song/album articles, then goes back to being idle. Nothing else about the units changed. The weekly inspiredby-refresh.timer (Sun 04:00) now also queues any new song/album articles from the dump.
- New sources: Wikipedia full-text search, Wikidata dedicatee / commemorates / named-after links, and "research leads" verified against Wikipedia (a trial of 120 written from Claude's memory got 74 verified, with both deliberately false controls rejected).
- Per-source yield is logged in a new
task_runstable;python3 run.py statsranks entry points by new claims per 100 requests. - Production DB backups taken before each bulk re-derivation:
~/inspiredby/backups/inspiredby-2026-10-05-pre-reextract.db,…-pre-dump.db. - The bot also fetches a cited web page once when a research lead quotes it, to check the quote. Same User-Agent, cache and throttle as before.
3. Docs catch-up
inspiredby (port 8440, its NPM host, its timers) had never made it into the cheat sheet or the context dump, so I added it there along with the new data folder.
Net effect: Wikipedia lookups on this box are now local and unlimited, and inspiredby is back to gathering, at CPU speed rather than Wikimedia's rate limit.
← Back to Admin Hub