SKYHOUSE.dev Journal

Maintaining the Cloud Fortress

The PID That Outlived Its Process

Why: Electrical work in the server room meant a planned sequence of soft shutdowns, but one power cut arrived unannounced — and Plex did not come back afterwards.

Four boots in one afternoon. Three of them were clean, deliberate shutdowns taken around the electrician's schedule. The fourth was not: the boot that started at 18:48:53 simply stops mid-line in the journal at 18:49:37, forty-four seconds later, with zero shutdown markers. That is the unplanned cut. Everything below follows from it.

1. The symptom: Plex down for an hour, retrying and failing

After the last boot, plexmediaserver.service was in failed state and had been for an hour. The journal showed the same line five times:

Plex Media Server is already running. Will not start...
plexmediaserver.service: Main process exited, code=exited, status=1/FAILURE
...
plexmediaserver.service: Start request repeated too quickly.

Nothing was listening on 32400, and no Plex process existed. The server was insisting it was already running while demonstrably not running.

2. The mechanism: a PID file that got reused out from under it

Plex writes its process ID to plexmediaserver.pid in the application support directory and removes it on a clean exit. The 44-second boot never got a clean exit — the power went out with the file still on disk, holding PID 1324.

On the next boot, PID 1324 was handed out again, this time to (sd-pam), an ordinary systemd session helper. So when Plex started, it read the stale file, checked whether that PID was alive, found that it was, and concluded a second copy of itself was already running. The check is sound; the assumption that a live PID means its own live PID is not.

The five retries then made it worse rather than better. systemd's StartLimitBurst tripped after five attempts in fifty seconds and gave up entirely, which is why the service stayed down for an hour instead of eventually recovering on its own.

3. The fix

Verified active, listening on 32400, and /identity returning HTTP 200. Total fix time was under a minute once the cause was clear.

4. What we checked next, and why

A hard power cut with Plex mid-write is the classic way to corrupt the library database, so that was the real thing to worry about — the service failure was only noisy. PRAGMA integrity_check against the 548 MB com.plexapp.plugins.library.db returned ok, with no malformed-image or SQLITE_CORRUPT entries in the logs. The database came through intact.

5. The film that wasn't missing

A copy of The Secret of NIMH appeared to have vanished from /media/plex1/Movies. It had not. The file was present and byte-identical to its mirror on plex1_backup, read cleanly at start, middle and end, and was still in the library database with an empty deleted_at.

The explanation was two overlapping blackouts. /media/plex1 was unmounted from roughly 13:52 to 18:49 — the DAS is a deliberate post-boot hotplug and had not been re-attached, so systemd retried media-plex1.mount every few minutes and failed each time. During that window every plex1-backed library was dark, all 1,606 films, not one. Then Plex itself was down from 18:51 to 19:53 with the PID problem above. Depending on when you looked, you saw either an empty library or no server.

The genuinely dangerous version of this did not happen: a Plex library scan against an unmounted drive can mark every item as missing and gut the library. No library scan ran in that window — only plug-in scanning at startup — and all 6,101 plex1 media parts were intact afterwards, with nothing marked deleted.

6. What this changes

One unplanned power cut, one hour of Plex downtime, five hours of dark libraries — and no data lost anywhere.

← Back to Admin Hub