๐ Day 191 โ Building a Watchdog for Expiring Credentials and Proving the Alarm Fires
๐ Topic
The personal secretary went quiet, and the root cause traced back to Google OAuth credentials expiring with no alert firing. I built an independent watchdog โ then made myself prove the alarm actually rings before trusting it.
๐ฏ Goal
Stop silent OAuth expiry from taking down a system nobody is watching in real time, and make sure the watchdog built to catch it is itself proven to work, not just installed.
๐ What I Did
I treated the outage as a monitoring failure first, and a credentials problem second.
Main areas covered:
- diagnosed why the secretary died silently: a Google grant expired and nothing downstream noticed until real functionality broke
- built an independent Google credential watchdog with loud alerts, deliberately separate from the pipeline whose health it reports on
- delivered a one-time verdict when a grant outlives its 7-day cap, so an aging credential gets exactly one clear warning instead of silence followed by failure
- confirmed the auth watchdogโs expiry cause with direct evidence rather than assuming the diagnosis was right
- did not stop at โthe watchdog existsโ โ deliberately forced the alarm condition and confirmed it actually fires, loudly, where I would see it
- shipped a self-contained, copy-pasteable reauth command so recovery does not depend on remembering the right incantation at 11pm
๐ Key Cybersecurity Connections
A monitor that has never been proven to fire is a monitor you believe in, not a monitor that works. This is the same lesson as the hermes-verify false-positive investigation and the fail-closed supervisor work โ except inverted: instead of a check crying wolf, this was a check that stayed silent when it should have screamed. Silent failure of a monitor is worse than a noisy false positive, because nothing draws attention to it.
Building the watchdog independently of the system it watches matters too โ a health check that shares a failure mode with the thing it monitors can go dark at exactly the moment itโs needed.
๐ Investigation Questions
- Was the credential expiry itself the root cause, or a symptom of something not watching for it?
- Does the watchdog share any dependency with the system it monitors?
- Has the alarm condition ever actually been triggered and observed, or only coded and assumed?
- Is recovery documented as a copy-pasteable step, or as tribal knowledge?
- What is the blast radius of one more silent expiry before this fix?
๐จ Detection Opportunities
Checks for a credential-health watchdog:
- a grant approaching or past expiry with no alert raised
- watchdog logic sharing infrastructure with the system it monitors
- an alert path that has never been deliberately tested
- recovery steps that exist only in memory, not in a runnable command
- repeated silent failures of the same category before the first real fix
Example:
project=google-auth-watchdog
signal=credential_expiry_with_no_alert_raised
risk_area=unmonitored_dependency
triage=confirm_watchdog_independence_and_force_test_the_alarm
๐งญ MITRE ATT&CK Techniques
No direct mapping claimed. This is availability monitoring and operational-resilience work for a credential dependency.
๐บ Visual Investigation Diagram
Google grant expires
โ
Old: nothing notices until real failure
โ
Independent watchdog built
โ
Deliberately force the alarm condition
โ
Confirm it actually fires, loudly
โ
One-time verdict + copy-pasteable reauth
โ Challenges
The uncomfortable question was โhow do I know this watchdog works?โ โ and the honest answer, before today, was that I didnโt. Forcing the condition felt like extra work for something that โshould obviously work.โ It wasnโt extra work; it was the actual point.
๐ What I Learned
I learned that โI built a monitorโ and โI proved the monitor firesโ are different claims, and only the second one is worth anything during an actual incident.
โก Next Steps
- Schedule a periodic forced test of the watchdog, not just a one-time proof
- Extend the same independence principle to other credential-backed integrations
- Keep the one-time-verdict pattern for other aging-grant scenarios
- Document the forced-test procedure so it survives me forgetting it exists
๐ง Reflection
An alarm that has never gone off is a hypothesis, not a safeguard. Making it ring on purpose, before I needed it to, is what turned this from a plausible fix into a trusted one.
๐งฉ Lessons Learned
What worked
Building the watchdog independently of the system it monitors, then deliberately forcing the alarm to prove it fires.
What broke
Google OAuth expired silently and took down the personal secretary with no warning.
Why it broke
Nothing was watching the credentialโs health independently of the pipeline that depended on it.
Fix / takeaway
Monitors need proof of firing, not just proof of existing โ and they should never share a failure mode with what they watch.
๐ Skill Progression Context
This supports my cybersecurity progression because credential lifecycle monitoring, independent health checks, and proving alerts actually fire are foundational SOC and reliability-engineering skills.
๐ TL;DR
Built an alarm for expiring credentials, then made myself prove it actually rings before trusting it.
