Server Cheat Sheet

Reference card for skyhouse.dev

The "what is what / where is what" reference for this server. For "something is wrong, what do I do" see the Troubleshooting Guide; for full prose orientation see /home/plex/docs/server-context.md. Verified 2026-05-20.

1. Network & access

ItemValue
Host / LAN IPplex — 192.168.1.136 (static DHCP reservation)
SSHport 2222 (not 22) — ssh -p 2222 plex@skyhouse.dev
TopologyISP modem (bridge mode) → GL.iNet router (192.168.1.1) → server
RouterGL.iNet BE9300 (Flint 3) — Qualcomm IPQ5332, tri-band Wi-Fi 7, fw 4.9.0 (OpenWrt 23.05-SNAPSHOT). Cheat sheet said v4.8.4 until 2026-08-18. Root SSH from server: ssh -i ~/.ssh/router_ed25519 root@192.168.1.1. Full docs: /home/plex/docs/router-reference.md — read before debugging any Wi-Fi issue. Wi-Fi driver is Qualcomm proprietary: iw ... station dump returns NOTHING, use wlanconfig <if> list sta; iw ... survey dump does work. Port-forwards live in /etc/config/port_forward, not firewall. SSIDs: SH24 (2.4/ch1), SkyHouse+SkyHouseWork5 (5G/ch36/EHT160), SkyHouse6+SkyHouseWork6 (6G/ch5/EHT320). Band steering + MLO + SQM all off. Idles at 71°C against a 75°C fan trigger. [Journal: 2026-08-18 wifi-rate-collapse]
Router port-forwardsWAN 80, 443, 32400, 2222 → 192.168.1.136 (same ports)
Emergency hostnamecw526dc.glddns.com — GL.iNet DDNS, always tracks current WAN IP
DNSCloudflare anycast NS (luciane/remy.ns.cloudflare.com) for skyhouse.dev, botaa.org, room101.com (registered at Dynadot). All records DNS-only (grey cloud). Apex = flattened CNAME → cw526dc.glddns.com (router DDNS) on all three — genuinely done 2026-08-04, so IP changes self-heal with no scripts and no hardcoded origin IP; * wildcard → apex. (Before 2026-08-04 all three were static A records despite docs claiming otherwise — room101 served a dead IP for 16 days.) room101 email = Cloudflare Email Routing. dns-monitor.sh notifies only. [Journal: 2026-08-04 apex-cname-flattening]

2. Port map

Red = forwarded to the internet. Green = LAN / localhost only.

PortServiceScope
80 / 443Nginx Proxy Manager — HTTP / HTTPS for every subdomainPublic
81NPM admin UI — manages all proxy hosts & certsLAN only (UFW-restricted)
32400Plex Media Server (main)Public (Plex remote access)
32401Plex (local control)localhost
2222SSH / SFTPPublic
9443 / 9000Portainer — the web UI that manages all Docker containers (9443 = HTTPS, preferred)LAN only
2283Immich (photos)LAN; public via photos.skyhouse.dev
3000Pairdrop (file sharing)LAN; public via share.skyhouse.dev
5001SharedmomentsLAN; public via btst.skyhouse.dev
8000 / 8001Music Discovery — client / API serverLAN; public via music.skyhouse.dev
3001Uptime Kuma (status page)LAN; public via status.skyhouse.dev
19999Netdata (system monitoring)LAN; public via monitor.skyhouse.dev
8989Sonarr (TV PVR)LAN only (external access not yet wired — see 2026-05-28 journal)
7878Radarr (Movie PVR)LAN only (external access not yet wired)
8686Lidarr (Music PVR)Public via lidarr.skyhouse.dev (basic auth, bypassed for LAN IPs) — 2026-07-17 journal
6969Whisparr (V3/Eros, adult movie PVR)Public via whisparr.skyhouse.dev (basic auth, bypassed for LAN IPs) — 2026-07-17 journal
6500rdt-client (Real-Debrid client)LAN only
8123Home Assistant (host-networked)LAN; public via ha.skyhouse.dev
8095Music Assistant (host-networked)LAN; public via ma.skyhouse.dev
8440inspiredby website (systemd user service inspiredby-web, read-only on its DB)LAN + ufw 172.18.0.0/16 (NPM); public via inspiredby.skyhouse.dev [journal]
8430yt-plex (systemd user service, uvicorn)LAN + ufw 172.18.0.0/16 (NPM); public via yt.skyhouse.dev [journal]
9876ddns-go UI — legacy, slated for retirementLAN only
53systemd-resolved (local DNS stub)localhost
internalVaultwarden :80, immich Postgres :5432, immich Redis :6379, npm-db MariaDB :3306, npm-php :9000Docker network only (no host port)

3. Reverse-proxy (subdomain) map

Authoritative hostname → backend, as configured in Nginx Proxy Manager.

HostnameBackend
skyhouse.devstatic site — /home/plex/www (legacy sites served as path subfolders, e.g. /randoma); *.php → npm-php-1 (php-fpm). PHP location needs its own root /var/www/html; in host 9 Advanced config, fixed 2026-10-08. Fire 7 kiosk control at /serverjournal/kiosk/ (password; tablet polls config.php every 60 s) [journal]
plex.skyhouse.dev192.168.1.136:32400
photos.skyhouse.dev192.168.1.136:2283 (Immich)
music.skyhouse.dev192.168.1.136:8000 (/api → :8001)
vault.skyhouse.devvaultwarden:80 — LAN-allowlisted, not public
share.skyhouse.dev192.168.1.136:3000 (Pairdrop)
btst.skyhouse.dev192.168.1.136:5001 (Sharedmoments)
monitor.skyhouse.dev192.168.1.136:19999 (Netdata)
status.skyhouse.dev192.168.1.136:3001 (Uptime Kuma)
router.skyhouse.devhttps://192.168.1.1 (GL.iNet admin)
portainer.skyhouse.devhttps://192.168.1.136:9443 (Portainer)
npm.skyhouse.devhttp://127.0.0.1:81 (NPM admin โ€” LAN only)
sonarr.skyhouse.dev192.168.1.136:8989 (Sonarr)
radarr.skyhouse.dev192.168.1.136:7878 (Radarr)
rdtclient-tv.skyhouse.dev192.168.1.136:6500 (RDTClient TV — cache-only, for Sonarr)
rdtclient-movies.skyhouse.dev192.168.1.136:6501 (RDTClient Movies — open slots, for Radarr)
prowlarr.skyhouse.dev192.168.1.136:9696 (Prowlarr — indexer manager)
lidarr.skyhouse.dev192.168.1.136:8686 (Lidarr — Music PVR, library at /media/plex3/Music)
whisparr.skyhouse.dev192.168.1.136:6969 (Whisparr V3/Eros — adult movie PVR, library at /media/plex1)
ha.skyhouse.dev192.168.1.136:8123 (Home Assistant — host IP, not a container name; WebSockets on)
ma.skyhouse.dev192.168.1.136:8095 (Music Assistant — host IP; WebSockets on + hardened Advanced location block for its auth handshake)
inspiredby.skyhouse.dev192.168.1.136:8440 (inspiredby — public, read-only song-inspiration index). [journal]
yt.skyhouse.dev192.168.1.136:8430 (yt-plex — app does its own login; live progress is SSE, which the app un-buffers with X-Accel-Buffering: no). Live 2026-09-29. [journal]
yt-plex.skyhouse.devyt-plex-demo:8430 (Docker DNS on npm_default; public yt-plex demo, no login, sandbox per visitor; Let's Encrypt + Force SSL). Set up by Damien 2026-10-01. [journal]
room101.comstatic — /home/plex/www/room101
botaa.org / share.botaa.org / smut.botaa.orgstatic / Pairdrop / Plex (group connect URL)

4. Service inventory

Docker containers (23)

ContainerPurposeCompose dir
npm-app-1 / npm-db-1 / npm-php-1Nginx Proxy Manager + MariaDB + PHP-FPM/home/plex/npm
immich_server / _postgres / _machine_learning / _redisImmich photo platform (4-container stack)/home/plex/immich-app
music-discovery-client / -serverNext.js music app + API/home/plex/music_discovery
uptime-kumaService uptime monitoring/home/plex/uptime-kuma
yt-plex-demoPublic yt-plex demo (YTP_DEMO=1): real lookups, nothing downloaded/saved; per-visitor SQLite sandboxes in /home/plex/yt-plex-demo/data; Radarr/Sonarr keys in /home/plex/yt-plex-demo/demo.env (600). Networks npm_default + media-downloads_default, no host port, read-only rootfs, no media mounts. Image yt-plex-demo:latest (rebuild from /home/plex/yt-plex, then recreate with the docker run in the journal) [journal]standalone docker run
vaultwardenPassword vault (Bitwarden-compatible)standalone docker run
pairdropCross-device file sharingstandalone
sharedmomentsShared photo albumstandalone (network btst_default)
portainerDocker management web UI — pinned to portainer-ee:2.42.0 after May 27 :latest drift (journal)standalone
prowlarrIndexer manager for Sonarr & Radarrportainer stack, port 9696
lidarrMusic PVR — library at /media/plex3/Music, config at /home/lidarr/data (journal)portainer stack, port 8686
whisparrโ›” PARKED 2026-08-27 — stopped, service commented out of the Portainer stack, removed from das-up.sh and das-mount-check.py, and NPM proxy host 39 (whisparr.skyhouse.dev) disabled. It was never configured (0 movies, 0 indexers, 0 download clients) yet ran RSS syncs every 30 min, sat publicly on 443, and held read-write access to all of /media/plex1 — and had been bound to a stale empty DAS mount for as long as das-up.sh existed, unnoticed because nobody uses it. Nothing was lost: config is a host bind mount at /home/whisparr/data, so root folder and settings survive. To revive: uncomment in the stack → redeploy → re-add to both scripts → re-enable proxy host 39 (backups: /home/plex/ops/whisparr-parked-20260827/). Keep restart="no". [journal] Adult movie PVR, V3/Eros build (ghcr.io/hotio/whisparr:v3 — different env convention than the linuxserver images: UMASK/WEBUI_PORTS) — library at /media/plex1 (whole drive, root folders set inside its own UI, same as Radarr), config at /home/whisparr/data (journal)portainer stack, port 6969
sonarrTV PVR — library at /media/plex2/TV, config at /home/sonarr/data. After the split-rdtclient migration, root folder inside container = /data/TV (journal)portainer stack, port 8989
radarrneeds /media/plex1 (hotplug DAS); restart=no — started via das-up.sh after attaching the DAS. Movie PVR — library at /media/plex1/Movies, config at /home/radarr/data. After split, root folder = /data/Movies; download client reaches rdt-client by Docker service name on the internal port (rdtclient-movies:6500, not host-IP:6501) (journal, rewiring, plex1 offline)portainer stack, port 7878
rdtclientRDTClient TV (for Sonarr; intended cache-only but not actually gated — see 2026-05-29 journal) — downloads to /media/plex2/torbox_downloads, db at /home/rdtclient/db (journal)portainer stack, port 6500
rdtclient-moviesneeds /media/plex1 (hotplug DAS); restart=no — started via das-up.sh after attaching the DAS. RDTClient Movies (open slots, for Radarr) — downloads to /media/plex1/torbox_downloads, db at /home/plex/rdtclient-movies/db. Compose at /home/plex/rdtclient-movies/docker-compose.yml (journal, plex1 offline)portainer stack media-downloads (id 46), port 6501 — corrected 2026-08-27; previously recorded here as standalone compose
homeassistantHome Assistant — network_mode: host, port 8123. Config at /opt/appdata/homeassistant/config. HubZ Zigbee/Z-Wave stick passed through via two /dev/serial/by-id paths → ttyUSB0/ttyUSB1. Also mounts /etc/localtime + /run/dbus. [journal]portainer stack home-assistant, port 8123
music-assistantMusic Assistant — network_mode: host, port 8095. Data at host /home/music-assistant → container /data; library /media/plex3/Music mounted ro. Caps: SYS_ADMIN, DAC_READ_SEARCH, apparmor:unconfined. [journal]portainer stack home-assistant, port 8095
ddns-goLegacy DDNS updater — slated for retirementstandalone

systemd services (non-stock)

plexmediaserver · netdata · fail2ban · smartmontools · ssh (port 2222) · docker / containerd. User units (systemctl --user, linger on): inspiredby-web (port 8440) + inspiredby-worker.timer (hourly; drains its task queue for up to 50 min, budget = 250 network-using tasks) + inspiredby-refresh.timer (Sun 04:00) + inspiredby-reader.timer (hourly; Claude Haiku via claude -p under Damien's login reads 100 Wikipedia articles/h; INSPIREDBY_LLM_READER=off / _PER_HOUR in .env) [journal] — PAUSED 2026-10-07 (INSPIREDBY_LLM_READER=off, timer stopped: ~1M tokens/100 pages on a Pro plan); hosts that send a long Retry-After are blocked in ~/inspiredby/.cache/blocked_hosts.json (delete an entry to lift it early) [journal]; Spotify matching capped at 1 task/run, 10/day, plus site play buttons that look songs up on demand (/api/spotify, 120/h site-wide); Brave Search counter in ~/inspiredby/.cache/brave_usage.json (run.py brave --remaining N to correct); overnight reader test scheduler/usage-test.sh full [journal], code /home/plex/inspiredby — secrets in ~/inspiredby/.env (600: Last.fm, Spotify client shared with music-discovery, cookie key); public POST /suggest writes only to ~/inspiredby/submissions/; photos in ~/inspiredby/media/photos/; DB is WAL (back up with sqlite3 .backup) — admin mode: direct LAN http://192.168.1.136:8440/admin, editor-pass cookie, or editor Spotify id (never proxied IPs: NPM trusts X-Real-IP from CDN ranges) [journal] [journal] (deployed from ~/dev/inspiredby via scheduler/deploy.sh; start with its CLAUDE.md) [journal]; yt-plex — yt-dlp → Plex queue app, port 8430, code at /home/plex/yt-plex. v2 (2026-09-29): login, libraries and Plex/Radarr/Sonarr connections live in the app DB (data/ytplex.db, edit via Settings); no .env; yt-dlp self-updates daily in-app into data/ytdlp-lib (the old 50 5 * * * pip cron is retired). Also packaged as a Docker image (yt-plex:latest, built locally). Published at skyhouse.dev/yt-plex/ (/home/plex/www/yt-plex/). Lost password: YTP_RESET_PASSWORD=โ€ฆ in data/.env + restart. [journal] [journal v2] [journal] (openvpn was disabled in the 2026-05 cleanup.) Note: plexmediaserver is temporarily masked as of 2026-06-04 (plex1 SATA link wedged); systemctl unmask plexmediaserver to restore. [journal]

5. Key paths

WhatPath
Web root/home/plex/www
Server journal & these docs/home/plex/www/serverjournal/
Plex config / library DB/var/lib/plexmediaserver/Library/Application Support/Plex Media Server/
Docker data root/var/lib/docker (on the OS disk)
NPM config & certs/home/plex/npm/data/, /home/plex/npm/letsencrypt/
Maintenance scripts/usr/local/bin/ and /home/plex/bin/ and /home/plex/*.sh
Logs/var/log/plex-backup.log, ~/.cache/dns-monitor.log, ~/.cache/nextjs-check.log
Secrets (locations only)/etc/telegram_notify.conf, ~/.messaging-keys, /etc/netdata/health_alarm_notify.conf, /home/plex/npm/data/.htpasswd
Local Wikipedia (shared, read-only for projects)/home/plex/data/wikipedia/ — enwiki text dump + SQLite index (current/wikidump.db), library wikidump.py, see its README. ~34 GB on the OS disk. [journal]
Orientation docs/home/plex/docs/ — server-context.md, server-audit-2026-05.md, server-runbook.md, disk-cutover-guide.md
Audit archive (movable junk)/home/plex/archive/2026-05-cleanup/

6. Cron schedule (user plex)

โš  Jobs below are shown logically. As of 2026-08-04 every watcher actually runs as flock -n /run/lock/<name>.lock timeout <N> <script> — before that none of the six had a lock or a timeout, so a wedged Docker daemon left instances stacking every 5 min. Backups were already guarded internally. telegram_notify.sh also had no curl timeout and never checked its exit code (a failed alert looked identical to a delivered one); it now logs every send to /var/log/telegram_notify.log. Crontab backups: ~/.cache/plex.crontab.bak-20260804, /root/root.crontab.bak-20260804. [Journal: 2026-08-04 watcher-hardening]

WhenJobPurpose
*/5 * * * *cpu_watchdog.shHigh-CPU alerting
*/5 * * * *bin/dns-monitor.shCloudflare-era DNS notifier (๐ŸŒ IP-changed / โœ… resynced / โš ๏ธ); no DNS writes
*/5 * * * *bin/arr-pipeline-check.pyMedia pipeline probe โ€” stalled items in the Sonarr/Radarr queues, and new downloads, โ†’ Telegram. The only check that watches whether media actually moves. 2026-08-05: skips entirely while /run/lock/arr-stack-maintenance is set (the 03:30 config backup stops the stack on purpose), and needs 2 consecutive unreachable readings before reporting DOWN โ€” one blip is silent. 2026-08-21: moved */30โ†’*/5 and does its own alerting โ€” it sends โŒ/โš ๏ธ/โœ… via telegram_notify.sh, announces downloads that just started, and pushes Kuma always up so that monitor is a dead-man's switch only. TorBox lookup is now lazy (only on a run that found a stall). [journal] [noise fix] [alert rewrite]
6,16,26,36,46,56 * * * *bin/das-mount-check.py --applyStale DAS bind-mount guard (2026-08-27). Docker resolves a bind mount once, at container start, and creates a missing source as an empty directory. A plex1 container that starts before the DAS mount lands therefore binds that empty directory for its whole lifetime — and every external signal still looks healthy: container up, mountpoint real, files on disk. Radarr sat like that for two days emitting path does not exist or is not accessible for paths that plainly existed, which reads as a permissions fault and sends you looking in the wrong place. Compares the device number the container has for each DAS-backed destination against the device the host has for the source (via /proc/<pid>/mountinfo — no exec, so it works on images with no shell) and restarts any mismatch. docker start cannot fix this — it is a no-op on a running container, which is why das-up.sh reported success every time. Also starts the plex1 containers when the mount is healthy but they are down, which restart=no otherwise leaves stranded after a Docker daemon restart. [journal]
5 * * * *bin/rdtclient-cleanup.py --apply --include-arr-referenced --include-idlerdt-client stale-record cleanup (scheduled 2026-08-20; script written 2026-08-04 and unwired until then). THE MISSING HALF of the Radarr cleanup below — the arrs rebuild their queues from the download client, so an arr-layer delete alone just reappears on the next sync with the same queue id. Seven Joy.Ride.2021 records resurrected that way for 9 days while the hourly Radarr cleanup logged DID NOT STICK every run. Runs at :05, twelve minutes ahead of :17, so the client layer is clean before the arr layer re-syncs. --include-arr-referenced is required, not optional: without it the layers deadlock protecting each other. Both real gates remain — record must be errored and ≥24h old. โš  Now defers entirely while any arr is downloading or awaiting import, and caps at 20 deletions per run (2026-08-21): rdt-client's delete endpoint takes ~32 s per call, so an unbounded backfill spent half an hour hammering the service Radarr polls every minute for completed-download handling. Blind spot CLOSED 2026-08-21 by --include-idle, which also sweeps the two shapes it used to ignore — records rdt-client reports Finished, and records stalled with no seeders — once older than 3 days and absent from every arr queue in a healthy state. That last qualifier is the point: a stalled record is in the arr queue because the client holds it, so plain presence cannot count as proof of life. [journal] [2026-08-21]
17 * * * *bin/radarr-queue-cleanup.py --applyRadarr queue self-heal (2026-08-05) โ€” the counterpart to Sonarr's */5 cleanup, which had existed since 2026-05-29 while Radarr had none, so Radarr's monitor sat DOWN for days. Hourly (not */5): it is the only ops script that deletes. Offset to :17 so it never lands on the probe (:00/:30) or the 03:30 backup. [journal]
8,23,38,53 * * * *bin/torbox-slot-guard.py --applyTorBox slot guard (2026-08-21). A magnet with a dead swarm never fails on TorBox — it sits in checking with size=-1 and eta=8640000 (100 days), holding one of the three paid slots for ever. With slots full TorBox queues new adds; rdt-client reads “queued” as a failed add, retries, gets DIFF_ISSUE: Download already queued and parks a nameless dead record — so one dead magnet poisons every grab behind it. Evicts only when active + unfinished + 0 progress + 0 seeds + 0 peers + >30 min, and separately flushes orphaned add-queue entries. Four times an hour because a jam costs every download queued behind it. [journal]
12,42 * * * *bin/rdtclient-stuck-alert.pyTerminal-failure alert — the only media job that texts you (2026-08-21). Fires when retryCount ≥ TorrentRetryAttempts (rdt-client has stopped trying) and either TorBox still holds the release cached (recoverable by hand; nothing else will ever retry it) or the error is unrecognised. Everything the stack self-heals stays silent — the first draft fired on 17 records of which 16 were slot-jam residue. Runs inside the 180-min DeleteOnError window so there is time to act. Dedup state: ~/.cache/rdtclient-stuck-alert.json (--reset to clear). [journal]
45 5 * * *bin/media-stack-update.shMedia stack auto-update (2026-08-21). rdt-client has NO self-update — the banner only compares version strings against GitHub, and :latest is resolved at pull time, so a container serves its build for ever until re-pulled (ours ran v2.0.140 from 2026-07-17 while v2.0.142 shipped). Runs Watchtower --run-once --cleanup over the eight media-downloads containers by explicit name — immich, home-assistant, vaultwarden, npm and portainer are deliberately out of scope. Why a wrapper and not Watchtower's own scheduler: Docker creates a missing bind-mount source as an empty directory, so recreating radarr with the DAS detached would point it at a library of zero films and let it reconcile. The wrapper proves /media/plex1/Movies and /media/plex2/TV are mounted and hold >10 entries (mounted alone is not enough), and refuses while /run/lock/arr-stack-maintenance is held by the 03:30 backup. Silent unless something updated or was refused. Log: ~/.cache/media-stack-update.log. [journal] โš  FIXED 2026-09-02: the job had deferred on all 7 runs it ever made and never once completed. It polls each arr's queue and fails closed if a queue is unreadable; after whisparr was parked on 08-27 its API returned Connection refused every night, so every run deferred. Whisparr removed from the poll list and from CONTAINERS (which is passed to watchtower — leaving it there risked resurrecting a parked service). Fail-closed behaviour deliberately kept. [Journal: 2026-09-02]
40 4 * * *bin/torbox-cleanup.py --applyTorBox retention backstop (scheduled 2026-08-21; written 2026-08-04 and unwired until then — the same gap rdtclient-cleanup sat in). Deletes finished TorBox torrents >6 days old that no arr is tracking. With FinishedAction now removing on import this should usually find nothing. [journal]
*/15 * * * *bin/dns-apex-check.shApex correctness โ†’ Kuma push monitors (2026-08-04). Asserts each apex resolves via an external resolver to the same IP as cw526dc.glddns.com. Local resolution is useless here โ€” /etc/hosts split-horizon hid a 16-day room101 outage. [journal]
0 1 * * *backup-plex3.shโ›” DISABLED 2026-08-22 (RAM fault: no large writes to backup drives until resolved; re-enable after). [Journal: 2026-08-22 backup-cron] rsync plex3 → plex3_backup. RE-ENABLED + REWRITTEN 2026-07-30 (was commented out; old version passed a stale device node and would have mounted the My Passport as the target). Moved to 01:00 so it can't overlap the 03:30 arr-config job. [journal] โœ… 2026-09-02: --max-delete=1000 now present in /usr/local/sbin/backup-plex3-rsync.sh — the pre-condition for re-enabling this cron is met. rsync exits 25 rather than deleting past the ceiling, and the caller already reports non-zero as failure.
disabledbackup-plex2.shโš  No target drive exists — /media/plex2 (~11 TB) has NO COPY ANYWHERE. Script neutralised 2026-07-30 (it refuses to run) because it held a stale /dev/sdb reference, which is now the scratch SSD. Needs a ≥12 TB target or SnapRAID parity. [journal]
30 4 * * *backup-plex1.shโ›” DISABLED 2026-08-22 (RAM fault; re-enable after, and add --max-delete=1000 as a circuit breaker). [Journal: 2026-08-22 backup-cron] rsync plex1 → plex1_backup. RE-ENABLED 2026-07-30 — it had been commented out since June despite a crontab note claiming it was re-enabled on 2026-06-04, so plex1_backup held 2.3 TB against plex1's 12 TB. Script itself was already sound. First convergence takes ~3 nights (~10 TB at 6h/night). [journal] โœ… 2026-09-02: --max-delete=1000 now present in /usr/local/sbin/backup-plex1-rsync.sh — the pre-condition for re-enabling this cron is met. rsync exits 25 rather than deleting past the ceiling, and the caller already reports non-zero as failure.
30 3 * * *bin/arr-config-backup.sh*arr + rdt-client config snapshot (settings/API-keys/DBs) → /media/plex3_backup/config-backups/, keep 14. Stops the stack ~15s for a consistent SQLite copy. Restore with bin/arr-config-restore.sh. 2026-08-05: writes /run/lock/arr-stack-maintenance around the stop so the pipeline probe and both queue cleanups treat the outage as planned. The flag carries a 15-min deadline โ€” a killed backup can never mute monitoring for good. [journal] [flag]
0 4 * * *plex-checkPlex version alert. Logs to ~/.cache/plex-check.log. FIXED 2026-09-10: used to redirect to /var/log/plex-check.log, unwritable by plex — silently never ran. [journal]
0 5 * * *nextjs-checkNext.js version alert

7. Storage & drives

โš  2026-08-21 — /media/plex1 is mounted clean with errors and needs an offline e2fsck. 19 corrupt inodes + 31 bad block-bitmap groups, all confined to Immich thumbnails (regenerable); no originals affected — verified by SHA-1 against asset.checksum and by confirming zero originals were written during the corruption window. It has run silently for hours because its Errors behavior is Continue, so it never remounts read-only — check with sudo dumpe2fs -h /dev/<plex1> | grep -E "Filesystem state|FS Error count", not by waiting for a failure. Never fsck'd since the fs was created in 2022. The fsck is blocked on the RAM fault (ยง11) — rewriting metadata across 12.7 TB through bad memory is how a thumbnail problem becomes a catastrophe. Expect 15–45 min, not hours: only 2.2 M of 427 M inodes are in use. Add whisparr to das-down.sh first — it holds a plex1 bind mount and will block the unmount. [Journal: 2026-08-21] Scope established 2026-09-02: all 19 EBADMSG paths are directories under /media/plex1/Photos/thumbs/ — Immich thumbnails, i.e. regenerable derived data. No originals/movies/metadata affected. Damage is frozen: identical 31 block groups across 08-22 โ†’ 09-02 (8d20h uptime, zero new errors with the memmap= blacklist live). ext4 quarantines those groups, costing ~4.2 GB of capacity rather than corrupting writes.

DeviceMountRole
1 TB SSD "SSV8" (AA0โ€ฆ2477)/ + /boot/efiOS disk — replaced the ancient Corsair via live ddrescue clone 2026-06-03; root ext4 grown to 938 GB (UUID 879bcc42-โ€ฆ). [journal]
sdc1 (14 TB)/media/plex1Media — ~92% full
sdh1 (14 TB, ext4)/media/plex1_backupBackup of plex1 (USB ext.) — reformatted NTFS→ext4 2026-05-31, UUID 70d95e9e-โ€ฆ. [journal]
sdg1 (14 TB)/media/plex2Media
(removed)/media/plex2_backupDead drive retired — WD 9RHHSV3L (5,483 offline-uncorrectable, confirmed 2026-06-04 even via toaster) physically removed; backup-plex2.sh cron stays disabled (no target). [journal]
sdd1 (12 TB)/media/plex3Media. (No longer holds swapfile_extra — moved to the scratch SSD 2026-06-04.)
sde1 (10 TB)/media/plex3_backupBackup of plex3
Kingston 128 GB SSD (203S10Q4T73Z)/media/scratchScratch disk (internal SATA, ext4 117 GB, UUID 46b67b3d-โ€ฆ) — hosts Plex transcode temp (/media/scratch/transcode). Disposable: ~9.8 yr-old SSD, nothing irreplaceable here. nofail. [journal]

โš  Device letters shuffle — identify drives by serial/UUID, never /dev/sdX. The sdc/sdd/sde/sdg/sdh labels in the table above are historical and have already moved twice. As of 2026-07-30 the DAS (JMS567 bridge) runs at SuperSpeed with 5 populated bays — plex1 (ST14000NM001G/ZL2BG3W8), plex3 (WD120EDBZ/5QGWXSKF), plex3_backup (ST10000DM005/WP001CYM), plus an unassigned Toshiba 6 TB (HDWE160) and WD 2 TB (WD20EARS). The two WD easystores (plex2, plex1_backup) are separate USB externals. Health check: ~/das-check.sh (link speeds, max_sectors_kb, UAS binding, I/O pressure); attach/detach helpers remain ~/das-up.sh / ~/das-down.sh (das-down FIRST, then unplug). [Journal: 2026-07-30]

Swap: /swapfile 4 GB (OS disk) + /media/scratch/swapfile_extra 32 GB (scratch SSD, sw,nofail) — 35 GB total. Moved off plex3 → scratch 2026-06-04. Media drives mount by UUID in /etc/fstab with nofail. [Journal: 2026-06-04]
โš  2026-08-06: the scratch swap line had been commented out in /etc/fstab — the box was silently running on 4 GB, not 35 GB. Restored (backup /etc/fstab.bak-20260806-swap). Post-reboot check: swapon --show must list both files; if only /swapfile appears, the scratch entry has been disabled again. Note neither entry sets pri=, so after a reboot the root SSD is the preferred swap target and scratch is overflow only. [Journal: 2026-08-06]

SMART monitoring (rebuilt by-id 2026-06-04): /etc/smartd.conf lists every present drive explicitly by serial (/dev/disk/by-id/ata-<model>_<serial> -d sat, no DEVICESCAN) — boot SSD, scratch SSD, plex1, plex2, plex3, plex1_backup. The detachable toaster drive plex3_backup uses -d removable (auto-detects as SAT when present, and lets smartd survive it being unplugged). โš  smartd EXITS (status 16) if a non-removable listed device is absent — so when you retire or move a drive you MUST drop/adjust its by-id line, or smartd dies silently at next boot (exactly what a stale Corsair line did after the 6/3 boot swap). netdata's smartctl collector (/etc/netdata/go.d/smartctl.conf) now monitors everything present (device_selector: '* *', no /dev/sdX exclusion) — the JMicron bridge that once hung it is gone, and netdata only scans present devices, so it's shuffle-proof. Backups: /etc/smartd.conf.bak-20260604, /etc/netdata/go.d/smartctl.conf.bak-20260604. [Journal: 2026-06-04]

8. Maintenance scripts

ScriptDoes
warmclaude <time> [-h lead]Claude session preheater โ€” schedules a tiny Haiku ping (default 2h) before a planned work session so the 5h token block resets mid-session. -n dry-run, -l list. Transient systemd user timers (survive logout โ€” linger enabled 2026-08-05); log ~/.claude/warmclaude.log. Kit + /span-sessions skill: ~/claude-span-kit/. Published publicly (unlisted, no auth, no secrets) at https://skyhouse.dev/span-kit/ โ€” install anywhere with curl -fsSL https://skyhouse.dev/span-kit/install.sh | bash; re-publish after edits with ~/claude-span-kit/publish.sh (bump VERSION first) [journal] (2026-08-05)
sudo plex-updateDownload + install latest Plex Media Server
plex-check / nextjs-checkVersion checks — Telegram alert if behind
cpu_watchdog.shAlerts on sustained high CPU by non-allowlisted procs
telegram_notify.shShared Telegram-send helper used by the others. 2026-08-09: HTML-escapes the body by default โ€” it sends parse_mode=HTML, and a single <, > or & in arbitrary text makes Telegram reject the WHOLE message with a 400. That silently killed the rdt-client pre-wedge alert for days, because its own remediation snippet contained <id>. Callers composing deliberate markup set TELEGRAM_HTML=1 (only arr-remediation-ledger.py does). Also caps at 4000 chars and logs the API's error description to /var/log/telegram_notify.log. 2026-08-21: house style. The ๐Ÿ–ฅ๏ธ plex: header is gone (one server โ€” it never carried information and pushed the glyph off line one). Every message is now glyph + scope + object, then what is happening, then one โ†’ next step; the glyph answers only does this need you? [journal] [house style]
bin/arr-pipeline-check.pyMedia pipeline probe (cron */5) โ€” watches the FLOW: items that enter the arr queues and never leave. Classifies errored/no-start/no-progress/import-blocked, names the titles, reports only (never touches the queue). Since 2026-08-21 it also alerts directly (โŒ = you act, โš ๏ธ = being handled, โœ… = cleared or a download starting) and detects new downloads from its own state file โ€” a downloadId with no prior record IS a grab that just began, which is why this needed no webhook. Supports --dry-run. Keys: ~/.arr-keys.env (2026-08-04, rewritten 2026-08-21)
bin/ops-capture.pyIncident capture โ€” ops-capture.py <dns|backup|media|docker|disk|system> --reason "..." writes a bundle to ~/ops/incidents/ with output + runbook. Not web-served; bounded commands; refuses to run under 500 MB free (2026-08-04)
ops/runbooks/*.mdAnchored runbooks (media#stalled-queue, dns#apex-mismatch, backups#job-silent) โ€” referenced in alert text, copied into capture bundles (2026-08-04)
bin/dns-apex-check.shApex-vs-DDNS correctness check feeding Kuma push monitors; external resolver, dual-resolver fallback (2026-08-04)
bin/lib-kuma.shSourced helper: kuma_push <TOKEN_VAR> <up|down> <msg>. Returns 0 on every failure path โ€” must never fail the job it monitors (2026-08-04)
ops/provision-kuma.pyIdempotent Kuma monitor provisioning from code (--dry-run default; venv at ops/venv; creds ~/.kuma.env). Emits ops/push-tokens.env. 2026-08-05: now CONVERGES existing monitors (it used to treat "exists" as "correct", so a wrong setting could never be pushed out) and computes grace as heartbeat + retry ร— retries. โš  Kuma applies retryInterval only while retrying โ€” with maxretries=0 a declared grace period does not exist, which is what made every push monitor flap. [journal]
bin/immich-flickr-dedupe.pyFlickr size-variant de-duplication (2026-08-21, dry-run by default; --apply to retire). โš  Never filter on the _o suffix — it lies. A 1024px _b render re-uploaded to Flickr later earns a NEW photo id and exports as that upload's _o, so a 680x1024 copy and the true 4288x2848 original both end _o. Groups on the first 8+ digit Flickr id, ranks on pixels then bytes. Albums are the majority case: 426 of 783 retirable assets were in albums — the keeper joins the loser's albums BEFORE the loser is retired. Soft-delete to the 30-day trash; full CSV audit in ~/.cache/. First run: 783 retired, 764 album additions, 0 failures. Only sees Flickr-vs-Flickr; cross-source copies need bin/immich-nearmiss-dupes.py. [journal]
bin/arr-cached-search.pyCached-first release picker (2026-08-10, report-only; --apply to grab). TorBox's checkcached answers for arbitrary infohashes, so "is it instantly available?" is known BEFORE grabbing. Seeders are the wrong signal on debrid โ€” a 1-seeder cached release beats a 30-seeder uncached one. Fails closed if the cache check errors. [journal]
bin/arr-upgrade-scan.pyLibrary scanner (2026-08-10, report-only). Default: files a cached smaller copy could replace. --upgrade: sub-1080p titles worth improving (8 for ~3 GB, some negative cost). No --apply by design โ€” profile Any has upgradeAllowed=false and a smaller file is not an "upgrade", so replacement is delete-then-repick (--script emits the worklist). [journal]
bin/tv-duplicate-audit.pyTV duplicate audit (2026-08-05, read-only, on demand). Plex has duplicate detection for movie libraries only, so a duplicated episode is invisible in the UI. Finds SPLIT SERIES (one show under two folders โ€” Bad Sisters vs Bad Sisters (2022)), DUPLICATE EPISODES, and folders Sonarr cannot see. Marks which copy Sonarr TRACKS โ€” deleting the tracked one leaves the episode missing and a monitored series re-grabs it. Ignores Featurettes/Deleted Scenes (they carry episode codes; counting them reported 737 instead of 400). Baseline report: /home/plex/docs/tv-duplicate-audit-20260805.txt. [journal]
bin/arr-remediation-ledger.pyRemediation bookkeeping (2026-08-05). Both queue cleanups record every fix against the episode/movie (not the release โ€” blocklisting swaps releases, so a release-keyed count never sees the loop). Routine self-healing is silent; it texts only when the same title needs remediating 3 rounds in 7 days (โ‰ฅ4h apart, so one multi-release pass counts once), or when the cleanup call itself fails. arr-remediation-ledger.py report for the last 7 days. [journal]
bin/wifi-device-stress.sh
bin/cast-network-test.py
bin/run-arch-ab2.sh
bin/ma-cli.py
Wi-Fi / multi-room audio diagnostics (2026-08-18, on demand). wifi-device-stress.sh floods each wireless client in turn and reports its loss plus the collateral latency inflicted on others — use it to find devices that poison the medium (needs sudo). The cast scripts A/B individual streams against Cast groups using ~/.venv-cast (pychromecast); run-arch-ab2.sh is the one to use (interleaved + repeated, since the idle baseline drifts several ms). Full writeup: /home/plex/docs/multiroom-audio-reference.md. [journal]
bin/router-snapshot.shRouter health snapshot (2026-08-18, read-only, no cron — run on demand). SSHes to the BE9300 with ~/.ssh/router_ed25519 and dumps identity, thermals, WAN, per-radio airtime survey, per-station TX-rate vs RSSI, wireless/steering/MLO/SQM config, port-forwards and recent Wi-Fi log events to /home/plex/docs/router-snapshots/<date>.txt. First tool to run for any Wi-Fi complaint. [journal]
bin/dns-monitor.shCloudflare-era notifier: Telegram on IP-change + resync confirm; no DNS writes (2026-06-07). 2026-08-21: on the shared sender, body in main() (parse guard), a retraction arm so a warning raised outside the resync window can be closed (it could dangle for ever before), and a consecutive-failure counter that reports a WAN blackout after 30 min — a 35-minute outage went entirely unreported that day. [journal] Largely redundant since 2026-08-04 (apexes self-heal); its โš ๏ธ never names the failing domain and never repeats โ€” slated for replacement by an external check. A scoped Cloudflare API token now exists at ~/.cloudflare.env (600) for operator-initiated work only โ€” no cron or watchdog may use it.
bin/lib-pubip.shHardened multi-source public-IP fetch (used by dns-monitor)
bin/install-script.shโš  USE THIS TO EDIT ANY CRON SCRIPT. Writes a temp file beside the target, syntax-gates it (bash -n / py_compile + shebang + empty-file guard), inherits the mode, then rename(2)s it into place. Bash tracks a byte offset through a script rather than slurping it, so rewriting a live one in place lets a running shell resume at a stale offset and execute a mix of old and new code — which sent a false DNS alarm on 2026-08-21. Rename is atomic: a running shell keeps its old inode. install-script.sh <src> <dest> [journal]
bin/lib-rsync-report.shTurns rsync output into one actionable line โ€” how many files failed and which failed first. Sourced by backup-plex{1,3}.sh, which used to send tail -n 50 of the log as the alert body (2026-08-21)
Dynadot-era (RETIRED, rollback-only):Moved out of ~/bin on 2026-08-21 → ~/bin/retired/: ip-watch.sh, ip-recovery-verify.sh, dns-watchdog.sh, cf-migration-verify.sh, dns-push.sh, dynadot-update.py. None scheduled or sourced since 2026-06-07; they carried 8 dead Telegram templates. Restore steps in ~/bin/retired/README.md. Only useful if reverting NS to Dynadot. [journal]
backup-plex{1,3}.shNightly media mirroring. Both now run their rsync as root via fixed-arg wrappers (/usr/local/sbin/backup-plex{1,3}-rsync.sh, sudo NOPASSWD) so they can read root-owned trees (immich on plex1; pre-swap-backup-2026-06-02 on plex3). Both guard on mountpoint for both endpoints and never mount by device node. plex3 excludes config-backups so the mirror can't delete the arr snapshots. robust_rsync.sh is no longer used by any enabled job — it auto-mounts a device node passed as an argument, which is what made the old scripts dangerous. [journal]
bin/pre-swap-config-backup.shOne-shot (run manually) pre-boot-disk-swap safety backup: DB dumps (immich PG, NPM MariaDB, vaultwarden) + all configs/certs/compose//etc → /media/plex3/pre-swap-backup-<date>/. See swap runbook
bin/arr-config-backup.shNightly (03:30) consistent snapshot of the media-download stack's config — sonarr/radarr/prowlarr/rdtclient(-movies) settings, API keys, and SQLite DBs → /media/plex3_backup/config-backups/arr-config-<ts>.tar.gz (~40 MB, keep 14). Briefly stops the stack so WAL is checkpointed; refuses to run if the dest drive isn't mounted. [journal]
bin/arr-config-restore.shRestore the above after a settings loss. arr-config-restore.sh (no args) lists snapshots; arr-config-restore.sh latest or โ€ฆ <file> restores (prompts y/N). Saves a pre-restore-<ts>.tar.gz of current state first, so the restore is itself reversible. Then in Prowlarr: Settings → Apps → Sync App Indexers if indexers look off.
lock_ssh.sh / unlock_ssh.shDisable / enable SSH password auth (key-only toggle)

Updating apps, generally (2026-07-17): two different procedures depending on where the app lives.

9. Verification checklist

Copy-paste to confirm the server is healthy — e.g. after a reboot or the disk swap. Each line prints OK / a status when good.

Services

systemctl is-active plexmediaserver netdata fail2ban ssh smartmontools
for c in npm-app-1 immich_server music-discovery-client music-discovery-server \
         vaultwarden pairdrop sharedmoments portainer uptime-kuma ddns-go yt-plex-demo; do
  printf '%-26s ' "$c"; docker inspect -f '{{.State.Status}}' "$c" 2>/dev/null || echo MISSING
done
curl -sf http://localhost:2283/api/server/ping        # Immich -> {"res":"pong"}

Mounts & freshness

for m in / /media/plex1 /media/plex2 /media/plex3 /media/scratch \
         /media/plex1_backup /media/plex3_backup; do  # plex2_backup drive retired 6/4
  printf '%-24s ' "$m"; mountpoint -q "$m" && echo OK || echo "NOT MOUNTED"
done
find /var/log/plex-backup.log -mtime -2 | grep -q . && echo "backups fresh"
find ~/.cache/dns-monitor.log -mtime -1 | grep -q . && echo "dns notifier fresh"

External reachability

curl -s ifconfig.me ; echo                # current public IP
dig +short skyhouse.dev @1.1.1.1          # should match the line above

10. Accounts & externals

AccountDetail
DynadotRegistrar + DNS for skyhouse.dev, room101.com, botaa.org (the dsm-iii/dsm-3 domains exist; hosting retired)
GL.iNet router192.168.1.1 — firmware v4.8.4
Telegram bot@skyhouse_server_bot — all alert delivery. Credentials in the 3 secret files listed in §5; rotating means updating all three.
Plex.tvaccount damienmjones

11. Quirks & gotchas

← Back to Admin Hub