SKYHOUSE.dev Journal

Maintaining the Cloud Fortress

The Mount That Was There, Except Where It Mattered

Why: Radarr had been showing completed downloads stuck at 100% in its queue for two days, with an import error naming a path that plainly existed on disk — and the workaround had become "download it by hand", which does not travel.

The reported symptom was stalled downloads. It was not a stall. Every file had arrived: TorBox had them, rdt-client had written them into /media/plex1/torbox_downloads/radarr/, and they were sitting there complete. Radarr simply could not see them, and said so in the one way guaranteed to send you looking somewhere else:

That message offers you two explanations, and both are wrong. The path existed. The permissions were fine. This entry is about the third possibility the error never mentions.

1. A bind mount is resolved once, and then never again

Docker resolves a bind mount at container start, and if the source directory does not exist it helpfully creates it — empty. From that moment the container holds a reference to whatever it found. Mounting a real filesystem over that path on the host afterwards changes nothing for the container: it is still looking at the empty directory underneath, and it will keep looking at it until it is restarted.

Radarr had started 33 seconds before the DAS mount landed. So:

That asymmetry is what made it confusing from the outside. Downloads worked perfectly, because the component doing the downloading could see the disk. Only the component doing the importing was blind, and the two sit next to each other in the same stack looking at the same path.

The proof is a device-number comparison. The host had /media/plex1 on 8:97 (the DAS). Radarr's mount namespace had /data on 8:3 — the root filesystem. Same path, different disk, no error anywhere.

2. Every signal we had said the system was healthy

This is the part worth remembering. There was no monitoring gap in the usual sense — the monitors were working, they were just all looking at things that were genuinely fine:

Nothing in the stack compares what a container can see against what the host can see, so the one broken thing was the one thing nothing measured.

3. das-up.sh could never have repaired this, and reported success anyway

The recovery script for exactly this class of problem runs docker start $CONTAINERS. docker start is a no-op on a container that is already running. It does not re-resolve bind mounts. So on every DAS attach, das-up.sh ran, printed radarr    Up 33 seconds, and exited 0 — while radarr sat bound to nothing. That line was in the systemd log the whole time; it is the timestamp that gives the fault away, and it reads like a success.

Only docker restart re-resolves the mount.

4. The real cause: documentation said restart=no, the stack said unless-stopped

The plex1 containers are supposed to be restart=no. That is not a preference — it is a data-safety measure. If Radarr starts against an empty library it can reconcile against it and mark 1,600 films missing. das-up.sh says so in its own header comment, and the cheat sheet has said so since 2026-07-06.

All three plex1 arr containers were on restart=unless-stopped. The Portainer stack definition (media-downloads, stack 46) carried restart: unless-stopped for every service, so every redeploy of the stack silently reverted the policy and let Docker start them at boot ahead of the mount. The documented design and the deployed reality had been disagreeing for long enough that nobody was checking.

This was not unforeseen. The context dump had carried this note since 2026-08-21, six days earlier:

“Not changed, flagged for a decision: with unless-stopped they will auto-start at boot whether or not /media/plex1 is present, which is the behaviour the DAS operating model was written to prevent.”

That was right, and it was left as a decision rather than a change. What the note did not anticipate — what nobody would — is that starting without the DAS would not fail loudly. It would succeed, quietly, against an empty directory, and stay that way for as long as the container lived.

whisparr was worse: it bind-mounts /media/plex1 exactly like radarr, and it had never been in das-up.sh's container list at all. It was stale too, and had been silently for as long as the file has existed.

5. What changed

The checker was verified against a real reproduction, not just a dry run: a throwaway container was bound to a directory, a tmpfs was then mounted over that directory on the host, and the container was confirmed still reading the old one. The check flagged it, restarted it, and the container's view matched the host afterwards.

6. Whisparr, which nobody had ever set up, is now parked

Finding whisparr stale raised a better question than how to fix it: what was it doing at all? The answer was nothing, expensively.

That is a public attack surface and a 12 TB write handle in exchange for zero function. It is parked rather than deleted, because the config directory makes parking free:

Nothing was lost. The config lives at /home/whisparr/data as a host bind mount, so the root folder (/data/Pr0n), settings and database survive independently of the container — they would survive deleting it outright. Reviving is: uncomment the block, redeploy, add the name back to the two lists, re-enable the proxy host. Backups of the NPM row and vhost are in /home/plex/ops/whisparr-parked-20260827/.

Worth noting that the staleness check itself is not list-driven — it walks every running container. So if whisparr is ever started again, it is protected from this bug immediately, whether or not anyone remembers to re-add it to PLEX1_MANAGED.

7. A discovery on the way out: the “cleared” alerts are not true

Reconstructing when Damien was told about this fault turned up a separate and worse problem. Correlating /var/log/telegram_notify.log against Radarr's import history gives this for Hadestown:

The mechanism is in the probe: cleared fires when an item stops matching the stall criteria, and a cleanup job deleting the record satisfies that perfectly. rdtclient-cleanup runs at :05 and radarr-queue-cleanup at :17; the false clear landed at :20, on the first poll after both. So one of our own jobs gave up on the download, and the monitor reported that as success.

Then the measurement, which is the part that matters. Of the seven ✅ Radarr cleared messages ever sent, four name a title and can be checked against the 177 recorded import events. None of the four corresponded to an actual import. One imported five hours later, one twelve hours later, and two never imported at all. The other three use the aggregate form — 2 stalled items resolved — which drops the titles and therefore cannot be audited even in principle.

The matching is by normalised title prefix and the window is a generous ±30 minutes, both of which bias towards over-counting true clears. It still came out zero.

So the green checkmark — the one symbol whose entire job is to say “stop worrying” — has been wrong every time it has been used. That is a more expensive fault than the mount, because it does not just fail to inform, it actively misinforms, and four repetitions is enough to teach someone to ignore the channel. Nothing is changed here today: a separate work stream is picking this up, and the full reconstruction, the measurement method and the message inventory are handed off in /home/plex/projects/arr/05-messaging-handoff.md.

8. Left open deliberately

sonarr and rdtclient mount /media/plex2 and remain restart=unless-stopped, because there is no plex2-stack.service — setting them to no would leave them dead after a reboot with nothing to start them. They are covered by the ten-minute checker but not by the ordering guarantee that protects plex1. A plex2-stack.service bound to media-plex2.mount would close that, and is a deliberate not-yet rather than an oversight.

Two days of "stalled downloads" were one container looking at the wrong empty directory, and the fix that existed for it had been reporting success while doing nothing.

← Back to Admin Hub