π Day 209 β Fixing a Watchdog That Mistook a Network Failure for an Expired Credential
π Topic
Fixing an authentication watchdog that could turn a transient network failure into a confident but wrong credential-expiry diagnosis.
π― Goal
Make the Hermes Google-auth watchdog distinguish an actually revoked grant from temporary DNS, timeout, or offline failures.
π What I Did
The watchdog makes a small authenticated request for each Google account. Previously, any failed request could be recorded as a credential βdeath.β That was wrong: two accounts failed at the same instant and recovered without re-consent, which is evidence of a network interruption, not revoked refresh tokens.
I changed the watchdog to classify failures from the full error output:
revokedwhen authentication indicators such asinvalid_grantare present;unreachablefor timeout, DNS, or offline failures;okfor a successful check.
Only a real revocation can record a grant-death age or trigger the βpossible seven-day testing-mode expiryβ diagnosis. Unreachable checks now accumulate separately and remain quiet for short blips; after repeated failures, they produce a different message that says the credentials are probably fine and the network needs attention.
π Key Cybersecurity Connections
This is alert quality. A security control that detects every failure but diagnoses the wrong cause trains its owner to distrust itβor worse, to take unnecessary remediation steps.
π Investigation Questions
- What evidence distinguishes an auth failure from a connectivity failure?
- Is the decisive error text being read, or only the last line?
- Can a transient fault poison long-term state?
- Does each alert tell the operator the right next action?
- Have both failure classes been tested?
π¨ Detection Opportunities
Useful fields include:
failure_kind=revoked|unreachable- unreachable streak length
- grant age at confirmed revocation
- account scope and last successful refresh
Example:
event=google_auth_watchdog
account=personal
failure_kind=unreachable
consecutive_checks=3
action=check_network_or_dns; do_not_reauthorize_yet
π§ MITRE ATT&CK Techniques
Relevant defensive context:
- T1078 β Valid Accounts, where account health must be assessed accurately
- T1550 β Use Alternate Authentication Material, because refresh-token failures need careful interpretation
πΊ Visual Investigation Diagram
Authenticated health check fails
β
Scan full error evidence
βββ auth indicators β revoked β record age + actionable reauth alert
βββ timeout/DNS/offline β unreachable β count streak + network alert only if persistent
β Challenges
The original watchdog had a plausible story: Google grants can expire in testing configurations. That plausible story became dangerous when it was attached to the wrong kind of evidence.
π What I Learned
I learned that βcheck failedβ is an observation, not a diagnosis. Monitoring must preserve uncertainty until the evidence supports a specific cause.
β‘ Next Steps
- Keep separate tests for revoked and unreachable cases
- Review alert wording whenever failure classification changes
- Retain the latest successful check as context for incident triage
π§ Reflection
Good monitoring does more than produce alerts. It helps prevent the wrong fix from becoming the next incident.
π§© Lessons Learned
What worked
Classifying evidence before writing durable failure state.
What broke
Treating every failed network call as a dead credential.
Why it broke
Different failure modes shared one coarse result code.
Fix / takeaway
Record the cause category separately, and make the alert match that category.
π Skill Progression Context
This is practical incident-response thinking: verify what failed, preserve the evidence, and choose remediation that matches the failure.
π TL;DR
An offline check should tell me to inspect the networkβnot to re-authorize a credential that is still valid.
