The Switch Was the Firewall
My CI runner could not SSH to my production host. Same VLAN, both up, both healthy, and every packet between them disappeared. No firewall log. No error. No RST. Just silence.
My CI runner could not SSH to my production host. Same VLAN, both up, both healthy, and every packet between them disappeared. No firewall log. No error. No RST. Just silence.
The day I moved my bot's broker gateway from a tmux session to a proper systemd service, I created the most educational outage of the whole project, and the scariest part is how good everything looked while it was broken.
I built a backup vault for my lab: a small Rocky Linux container on the most restricted VLAN I have, the database zone. Fresh container, correct IP, correct gateway. First test:
Today I tagged a release of my trading bot, watched the pipeline march through lint, test, and build, all green, and then watched the deploy stage fail.
Here is an uncomfortable question for anyone running a homelab: if one of your servers died right now, how long would it take you to find out?