The Mount That Was There, Except Where It Mattered
Why: Radarr had been showing completed downloads stuck at 100% in its queue for two days, with an import error naming a path that plainly existed on disk — and the workaround had become "download it by hand", which does not travel.
The reported symptom was stalled downloads. It was not a stall. Every file had arrived: TorBox had them, rdt-client had written them into /media/plex1/torbox_downloads/radarr/, and they were sitting there complete. Radarr simply could not see them, and said so in the one way guaranteed to send you looking somewhere else:
Import failed, path does not exist or is not accessible by Radarr: /data/torbox_downloads/radarr/Hadestown The Musical (2026) [...]/. Ensure the path exists and the user running Radarr has the correct permissions
That message offers you two explanations, and both are wrong. The path existed. The permissions were fine. This entry is about the third possibility the error never mentions.
1. A bind mount is resolved once, and then never again
Docker resolves a bind mount at container start, and if the source directory does not exist it helpfully creates it — empty. From that moment the container holds a reference to whatever it found. Mounting a real filesystem over that path on the host afterwards changes nothing for the container: it is still looking at the empty directory underneath, and it will keep looking at it until it is restarted.
Radarr had started 33 seconds before the DAS mount landed. So:
- On the host,
/media/plex1was a real mountpoint with 22 entries and 12 TB of library. - Inside
radarr,/datawas an empty directory. Not a broken one — an empty one. rdtclient-movies, which started 33 seconds later, caught the mount and was completely healthy.
That asymmetry is what made it confusing from the outside. Downloads worked perfectly, because the component doing the downloading could see the disk. Only the component doing the importing was blind, and the two sit next to each other in the same stack looking at the same path.
The proof is a device-number comparison. The host had /media/plex1 on 8:97 (the DAS). Radarr's mount namespace had /data on 8:3 — the root filesystem. Same path, different disk, no error anywhere.
2. Every signal we had said the system was healthy
This is the part worth remembering. There was no monitoring gap in the usual sense — the monitors were working, they were just all looking at things that were genuinely fine:
docker ps— radarr healthy, up 2 days.mountpoint /media/plex1— true.- The files — on disk, correct size, correct place.
- rdt-client — reporting the downloads finished, because they were.
- The pipeline probe — watches the arr queues for items that never leave, and it did see them; the alert was accurate and we still read it as a stall.
Nothing in the stack compares what a container can see against what the host can see, so the one broken thing was the one thing nothing measured.
3. das-up.sh could never have repaired this, and reported success anyway
The recovery script for exactly this class of problem runs docker start $CONTAINERS. docker start is a no-op on a container that is already running. It does not re-resolve bind mounts. So on every DAS attach, das-up.sh ran, printed radarr Up 33 seconds, and exited 0 — while radarr sat bound to nothing. That line was in the systemd log the whole time; it is the timestamp that gives the fault away, and it reads like a success.
Only docker restart re-resolves the mount.
4. The real cause: documentation said restart=no, the stack said unless-stopped
The plex1 containers are supposed to be restart=no. That is not a preference — it is a data-safety measure. If Radarr starts against an empty library it can reconcile against it and mark 1,600 films missing. das-up.sh says so in its own header comment, and the cheat sheet has said so since 2026-07-06.
All three plex1 arr containers were on restart=unless-stopped. The Portainer stack definition (media-downloads, stack 46) carried restart: unless-stopped for every service, so every redeploy of the stack silently reverted the policy and let Docker start them at boot ahead of the mount. The documented design and the deployed reality had been disagreeing for long enough that nobody was checking.
This was not unforeseen. The context dump had carried this note since 2026-08-21, six days earlier:
“Not changed, flagged for a decision: with
unless-stoppedthey will auto-start at boot whether or not/media/plex1is present, which is the behaviour the DAS operating model was written to prevent.”
That was right, and it was left as a decision rather than a change. What the note did not anticipate — what nobody would — is that starting without the DAS would not fail loudly. It would succeed, quietly, against an empty directory, and stay that way for as long as the container lived.
whisparr was worse: it bind-mounts /media/plex1 exactly like radarr, and it had never been in das-up.sh's container list at all. It was stale too, and had been silently for as long as the file has existed.
5. What changed
- Action: Restarted
radarrandwhisparr. Hadestown, the Emily Catalano special and a stranded Andrew Schulz release imported within a minute; the Radarr queue went to 0 and the inbox drained to empty. - Action: Set
restart=noonradarr,whisparrandrdtclient-moviesat runtime, and in both the Portainer stack file and the host-side canonical copy — otherwise the next stack redeploy would have quietly undone it. Note the value must be quoted asrestart: "no"; YAML parses a barenoas boolean false. - Action: Added
whisparrtoCONTAINERSindas-up.sh— then removed it again the same day and parked the service entirely. See §7. - Action:
das-up.shnow calls the new checker afterdocker start, so an already-running container bound to the wrong directory is caught on every DAS attach instead of being reported as fine. - Action: New
bin/das-mount-check.py, on cron at:06and every ten minutes after. For every container mount sourced under a DAS path it compares the device number the container has for the destination against the device the host has for the source, and restarts any container where they differ. It reads/proc/<pid>/mountinfofrom the host, so it needs noexec, no shell in the image and no writes to the media drives —portainerhas nosh, which rules out the obvious approach. - The same job covers the inverse:
restart=nomeans a Docker daemon restart leaves these containers down with nothing to bring them back, becausemedia-plex1.mountnever changed state andplex1-stack.servicetherefore never re-fires. If plex1 is mounted and a managed container is not running, it starts it.
The checker was verified against a real reproduction, not just a dry run: a throwaway container was bound to a directory, a tmpfs was then mounted over that directory on the host, and the container was confirmed still reading the old one. The check flagged it, restarted it, and the container's view matched the host afterwards.
6. Whisparr, which nobody had ever set up, is now parked
Finding whisparr stale raised a better question than how to fix it: what was it doing at all? The answer was nothing, expensively.
- 0 movies, 0 indexers, 0 download clients. It had never been configured.
- It ran an RSS sync every 30 minutes anyway, logging
No available indexers48 times a day. - It held read-write access to all 12 TB of
/media/plex1. - It was publicly reachable on 443 at
whisparr.skyhouse.dev. - And it had been bound to a stale, empty DAS mount for as long as
das-up.shhas existed — silently, precisely because nobody uses it. There was no user to notice.
That is a public attack surface and a 12 TB write handle in exchange for zero function. It is parked rather than deleted, because the config directory makes parking free:
- Action: Stopped the container; it is already
restart=no. - Action: Commented the whole service out of the Portainer stack (46) and the host-side copy, with the revival steps written into the comment block itself. Without this, the next stack redeploy would simply recreate and start it.
- Action: Removed it from
CONTAINERSindas-up.shand fromPLEX1_MANAGEDindas-mount-check.py, so neither tries to start a container the stack no longer defines. - Action: Disabled NPM proxy host 39 —
enabled=0in the database and the vhost renamed out of nginx's*.confinclude, then a validated reload.whisparr.skyhouse.devno longer answers;radarr.skyhouse.devwas checked as a control and still returns 200.
Nothing was lost. The config lives at /home/whisparr/data as a host bind mount, so the root folder (/data/Pr0n), settings and database survive independently of the container — they would survive deleting it outright. Reviving is: uncomment the block, redeploy, add the name back to the two lists, re-enable the proxy host. Backups of the NPM row and vhost are in /home/plex/ops/whisparr-parked-20260827/.
Worth noting that the staleness check itself is not list-driven — it walks every running container. So if whisparr is ever started again, it is protected from this bug immediately, whether or not anyone remembers to re-add it to PLEX1_MANAGED.
7. A discovery on the way out: the “cleared” alerts are not true
Reconstructing when Damien was told about this fault turned up a separate and worse problem. Correlating /var/log/telegram_notify.log against Radarr's import history gives this for Hadestown:
- 00:22 — Radarr grabs it.
- 02:25 —
❌ Radarr stalled — Hadestown. Correct, and consistently 2 h 03 m after the grab (ERROR_AGEplus one poll). - 03:20 —
✅ Radarr cleared — Hadestown. Not true. Nothing imported; Radarr's history has no event at all in that window, and it had to grab the film again nine hours later. - 15:16 — it finally imports, because I restarted the container.
The mechanism is in the probe: cleared fires when an item stops matching the stall criteria, and a cleanup job deleting the record satisfies that perfectly. rdtclient-cleanup runs at :05 and radarr-queue-cleanup at :17; the false clear landed at :20, on the first poll after both. So one of our own jobs gave up on the download, and the monitor reported that as success.
Then the measurement, which is the part that matters. Of the seven ✅ Radarr cleared messages ever sent, four name a title and can be checked against the 177 recorded import events. None of the four corresponded to an actual import. One imported five hours later, one twelve hours later, and two never imported at all. The other three use the aggregate form — 2 stalled items resolved — which drops the titles and therefore cannot be audited even in principle.
The matching is by normalised title prefix and the window is a generous ±30 minutes, both of which bias towards over-counting true clears. It still came out zero.
So the green checkmark — the one symbol whose entire job is to say “stop worrying” — has been wrong every time it has been used. That is a more expensive fault than the mount, because it does not just fail to inform, it actively misinforms, and four repetitions is enough to teach someone to ignore the channel. Nothing is changed here today: a separate work stream is picking this up, and the full reconstruction, the measurement method and the message inventory are handed off in /home/plex/projects/arr/05-messaging-handoff.md.
8. Left open deliberately
sonarr and rdtclient mount /media/plex2 and remain restart=unless-stopped, because there is no plex2-stack.service — setting them to no would leave them dead after a reboot with nothing to start them. They are covered by the ten-minute checker but not by the ordering guarantee that protects plex1. A plex2-stack.service bound to media-plex2.mount would close that, and is a deliberate not-yet rather than an oversight.
Two days of "stalled downloads" were one container looking at the wrong empty directory, and the fix that existed for it had been reporting success while doing nothing.
← Back to Admin Hub