This one isn't a session log. I went back through my own command reference — the second one, the topic-organized one — and what jumped out wasn't any single command. It was how often the fault hid behind a signal that looked fine.
So the note isn't the commands. It's the habit.
The wire can lie
The reference leads with diagnose bottom-up — link, then routing, then DNS, then the application — and the reason is an episode it keeps coming back to. One loose cable produced five different-looking failures: a broken git pull, name resolution failing, an unreachable router, an unreachable web UI, and an empty realm dropdown. Working it from the application layer down burned an hour. Starting at the wire would have found it in about two minutes.
The part I keep forgetting is that the link layer can negotiate while it's lying. A
partially-seated connector can hold enough contact to bring the link up and still drop
most of the frames. So an interface can carry a perfectly good address while the physical
link is dead — ip addr and ip link are different layers and they fail independently.
The same NO-CARRIER flag means "nothing on the other end" on a physical NIC and "no
containers attached" on a virtual bridge. Same word, opposite meaning. Don't chase the
second one.
When the software stack is demonstrably correct and nothing passes, suspect the wire. And when you need to split a routing fault from a DNS fault, the reference's trick is three ping targets: the router by address, a public address on the internet, and a hostname. If the public address answers but the hostname doesn't, the network is fine and only name resolution is broken.
A service that runs is not a service that works
systemctl status once showed active (running) while the messaging layer had been
deliberately suppressed. The process was up and serving HTTP; the channel was
intentionally off. Status told the truth about the unit and nothing about the product. The
reference's own rule is to verify by function — send a message, read a file back — not by
status.
The enable/disable trap is the same idea from the other side. start and stop act now;
enable and disable decide what happens at boot. A stopped service comes back on the
next reboot. A disabled one doesn't. And is-enabled can still miss it: a user-scope unit
carried over from an earlier host was starting on a headless box because lingering kept
that user's systemd manager alive from boot. The unit was already disabled. The door was
still open.
That one produced a real outage — two gateways competing for one fixed local port, eleven failed starts in five minutes, and a crash-loop breaker that suppressed the messaging channel. The thing worth remembering isn't a command; it's the enumeration habit: check system scope, check every user's user scope, and check lingering.
Narrow the fault by what still works
The cloud-storage section is where the habit paid off most clearly. Backups get pushed from the Proxmox host to Cloudflare R2 with rclone, and there were three separate 403s, each with a different cause. What broke the deadlock was noticing the asymmetry: listing the objects in one bucket succeeded while the copy failed. If the listing works, then the endpoint, the account, the bucket name and the credentials are all correct — which leaves only the permission model. Narrow a fault by what does work, not by staring at what doesn't.
The corollaries are the same instinct applied to access and to recovery:
- A 403 when a bucket-scoped token tries to list all buckets is least privilege working, not a bug. It was simply the wrong first test.
- Do not widen the token to make the error go away. Find the config flag.
- Read the file back. A backup you cannot retrieve is not a backup.
- The endpoint is account-level with no bucket appended. The URL on the bucket's settings page already contains the bucket, and using it doubles the path.
- Nightly rotating snapshots belong in the standard storage class. The infrequent-access class adds a retrieval fee and a thirty-day minimum storage duration, which is the wrong shape for something you delete after two days.
- Only the newest dump goes offsite, because the free tier is 10 GB and one snapshot is 4.8 GB compressed from a 32 GB disk, about a minute to create. Local keeps the history; offsite keeps the most recent copy.
- If a secret is ever exposed, delete the old token first and create the replacement second. That order leaves no window where the leaked key is still live.
The same rule for the plan I haven't run yet
The SSH hardening section is explicitly a plan, not something I've applied. The goal isn't
to remove root login — the Proxmox UI authenticates as a privileged realm account — it's
to stop password-based root login over SSH while keeping key-based access. The target
directive is PermitRootLogin prohibit-password, and specifically not the no form, which
breaks the Proxmox workflows.
The order is the safety mechanism, and it's the part I'd want my future self to slow down for:
- Generate the key.
- Push the public key, then confirm passwordless login from a second terminal before changing anything.
- Keep that working session and the physical console open as the way back in.
- Only then edit the config, restart the SSH service, and test from a new terminal.
Same rule as everything above: don't trust the change until something other than the thing you changed proves it.
What I'd actually want to remember
Every comfortable signal in this reference was wrong at least once. A link light up while
the cable was half-seated. A unit reading active (running) with its channel off. A unit
disabled while lingering quietly started it anyway. A bucket listing that succeeded while
the copy failed. A directive I'd planned but not yet applied.
The question that cut through all of them was the same: what still works, and what does that rule out? That's cheaper to keep than any command.