The site went down at 3 a.m.: how I find out first
Monitoring that wakes you only when it should. Two checks in a row, certificates via SNI, 80% disks and a false alarm during an update.
The worst way to learn your site is down is from a client. The second worst is a bot that screams “EVERYTHING IS ON FIRE” every five minutes when nothing is. After the third false alarm you mute it — and sleep peacefully through the fourth, real one .
I didn’t make up that second panel: it’s exactly what the very first dry run showed while I was updating the portal. Here’s how server alerts for bebb.be work now.
What to check
Every five minutes, without root access:
- Disks — over 80% is a reason to look, 100% is a reason to be sad.
- Services — nginx, databases, Redis, S3 storage, queues, the mail relay:
systemctl is-activefor each. - Certificates — not files, but what nginx actually serves.
- Queue — jobs that failed in the last hour.
- S3 backups — whether the latest upload of each policy succeeded.
Certificates via SNI, not from disk
You could read the date from the certificate file. But what usually breaks is something else: nginx serves the wrong certificate, hasn’t reloaded the new one, a domain ended up in the wrong vhost. So I ask nginx the way a browser does:
$ctx = stream_context_create(['ssl' => ['capture_peer_cert' => true, 'peer_name' => $host,
'SNI_enabled' => true, 'verify_peer' => false]]);
$s = stream_socket_client('ssl://127.0.0.1:443', $errno, $err, 8, STREAM_CLIENT_CONNECT, $ctx);
$cert = stream_context_get_params($s)['options']['ssl']['peer_certificate'];
$days = (openssl_x509_parse($cert)['validTo_time_t'] - time()) / 86400; // < 14 — alert
Two checks in a row
The main trick against false alarms: a problem must persist for two checks in a row (5–10 minutes) before anyone is notified. A service restart during an update takes seconds and goes unnoticed.
$st['seen']++;
if ($st['notified'] === null && $st['seen'] >= 2) { $new[] = $text; $st['notified'] = $now; }
elseif ($st['notified'] !== null && $now - $st['notified'] >= 6 * 3600) { $still[] = $text; } // reminder
Two more small things that save your nerves:
- “Back to normal” only for problems you were actually told about. Otherwise every restart would send “problem resolved” about a problem you never saw.
- Reminders every 6 hours, not every five minutes. An open problem won’t slip your mind, and your phone battery survives.
Where to send it
Two channels: email and Telegram. Email for the record, Telegram to read in a second. One message groups everything: “New problems”, “Still unresolved”, “Back to normal”. No twenty separate alerts in a row.
And client sites?
Client sites and bots have their own per-minute monitoring: the site fails to respond twice in a row, a bot is crash-looping, memory sits near the limit, a certificate expires within a week. The client is notified right away — by email and on Telegram — and the charts are in the portal .
Bottom line
- Check what the user sees (the certificate nginx serves), not what sits on disk.
- Two checks in a row — and false alarms all but disappear.
- Group alerts, remind rarely, report recovery.
- And sleep well. That’s what it was all for.