All notes

The site went down at 3 a.m.: how I find out first

Monitoring that wakes you only when it should. Two checks in a row, certificates via SNI, 80% disks and a false alarm during an update.

(-_-) zzZ 28 September 2026 · 3 min read DevOpsmonitoringLinux

The worst way to learn your site is down is from a client. The second worst is a bot that screams “EVERYTHING IS ON FIRE” every five minutes when nothing is. After the third false alarm you mute it — and sleep peacefully through the fourth, real one .

Dream monitoring
Speaks only when something breaks. One message. Clear.
Real-life monitoring
“Service reviactyl-queue: activating”. Because I just updated the panel and it’s restarting. Thanks, I know.

I didn’t make up that second panel: it’s exactly what the very first dry run showed while I was updating the portal. Here’s how server alerts for bebb.be work now.

What to check

Every five minutes, without root access:

  • Disks — over 80% is a reason to look, 100% is a reason to be sad.
  • Services — nginx, databases, Redis, S3 storage, queues, the mail relay: systemctl is-active for each.
  • Certificates — not files, but what nginx actually serves.
  • Queue — jobs that failed in the last hour.
  • S3 backups — whether the latest upload of each policy succeeded.

Certificates via SNI, not from disk

You could read the date from the certificate file. But what usually breaks is something else: nginx serves the wrong certificate, hasn’t reloaded the new one, a domain ended up in the wrong vhost. So I ask nginx the way a browser does:

$ctx = stream_context_create(['ssl' => ['capture_peer_cert' => true, 'peer_name' => $host,
    'SNI_enabled' => true, 'verify_peer' => false]]);
$s = stream_socket_client('ssl://127.0.0.1:443', $errno, $err, 8, STREAM_CLIENT_CONNECT, $ctx);
$cert = stream_context_get_params($s)['options']['ssl']['peer_certificate'];
$days = (openssl_x509_parse($cert)['validTo_time_t'] - time()) / 86400;   // < 14 — alert

Two checks in a row

The main trick against false alarms: a problem must persist for two checks in a row (5–10 minutes) before anyone is notified. A service restart during an update takes seconds and goes unnoticed.

$st['seen']++;
if ($st['notified'] === null && $st['seen'] >= 2) { $new[] = $text; $st['notified'] = $now; }
elseif ($st['notified'] !== null && $now - $st['notified'] >= 6 * 3600) { $still[] = $text; }  // reminder

Two more small things that save your nerves:

  • “Back to normal” only for problems you were actually told about. Otherwise every restart would send “problem resolved” about a problem you never saw.
  • Reminders every 6 hours, not every five minutes. An open problem won’t slip your mind, and your phone battery survives.

Where to send it

Two channels: email and Telegram. Email for the record, Telegram to read in a second. One message groups everything: “New problems”, “Still unresolved”, “Back to normal”. No twenty separate alerts in a row.

And client sites?

Client sites and bots have their own per-minute monitoring: the site fails to respond twice in a row, a bot is crash-looping, memory sits near the limit, a certificate expires within a week. The client is notified right away — by email and on Telegram — and the charts are in the portal .

Bottom line

  • Check what the user sees (the certificate nginx serves), not what sits on disk.
  • Two checks in a row — and false alarms all but disappear.
  • Group alerts, remind rarely, report recovery.
  • And sleep well. That’s what it was all for.

Got an idea? Let’s talk

Describe your task in a few sentences — I’ll reply with questions and a first estimate.