πŸ”„ Topic

Fixing an authentication watchdog that could turn a transient network failure into a confident but wrong credential-expiry diagnosis.


🎯 Goal

Make the Hermes Google-auth watchdog distinguish an actually revoked grant from temporary DNS, timeout, or offline failures.


πŸ›  What I Did

The watchdog makes a small authenticated request for each Google account. Previously, any failed request could be recorded as a credential β€œdeath.” That was wrong: two accounts failed at the same instant and recovered without re-consent, which is evidence of a network interruption, not revoked refresh tokens.

I changed the watchdog to classify failures from the full error output:

  • revoked when authentication indicators such as invalid_grant are present;
  • unreachable for timeout, DNS, or offline failures;
  • ok for a successful check.

Only a real revocation can record a grant-death age or trigger the β€œpossible seven-day testing-mode expiry” diagnosis. Unreachable checks now accumulate separately and remain quiet for short blips; after repeated failures, they produce a different message that says the credentials are probably fine and the network needs attention.


πŸ”— Key Cybersecurity Connections

This is alert quality. A security control that detects every failure but diagnoses the wrong cause trains its owner to distrust itβ€”or worse, to take unnecessary remediation steps.


πŸ” Investigation Questions

  • What evidence distinguishes an auth failure from a connectivity failure?
  • Is the decisive error text being read, or only the last line?
  • Can a transient fault poison long-term state?
  • Does each alert tell the operator the right next action?
  • Have both failure classes been tested?

🚨 Detection Opportunities

Useful fields include:

  • failure_kind=revoked|unreachable
  • unreachable streak length
  • grant age at confirmed revocation
  • account scope and last successful refresh

Example:

event=google_auth_watchdog
account=personal
failure_kind=unreachable
consecutive_checks=3
action=check_network_or_dns; do_not_reauthorize_yet

🧭 MITRE ATT&CK Techniques

Relevant defensive context:

  • T1078 β€” Valid Accounts, where account health must be assessed accurately
  • T1550 β€” Use Alternate Authentication Material, because refresh-token failures need careful interpretation

πŸ—Ί Visual Investigation Diagram

Authenticated health check fails
    ↓
Scan full error evidence
    β”œβ”€β”€ auth indicators β†’ revoked β†’ record age + actionable reauth alert
    └── timeout/DNS/offline β†’ unreachable β†’ count streak + network alert only if persistent

⚠ Challenges

The original watchdog had a plausible story: Google grants can expire in testing configurations. That plausible story became dangerous when it was attached to the wrong kind of evidence.


πŸ“š What I Learned

I learned that β€œcheck failed” is an observation, not a diagnosis. Monitoring must preserve uncertainty until the evidence supports a specific cause.


➑ Next Steps

  • Keep separate tests for revoked and unreachable cases
  • Review alert wording whenever failure classification changes
  • Retain the latest successful check as context for incident triage

🧠 Reflection

Good monitoring does more than produce alerts. It helps prevent the wrong fix from becoming the next incident.


🧩 Lessons Learned

What worked

Classifying evidence before writing durable failure state.

What broke

Treating every failed network call as a dead credential.

Why it broke

Different failure modes shared one coarse result code.

Fix / takeaway

Record the cause category separately, and make the alert match that category.


πŸ“ˆ Skill Progression Context

This is practical incident-response thinking: verify what failed, preserve the evidence, and choose remediation that matches the failure.


πŸ˜„ TL;DR

An offline check should tell me to inspect the networkβ€”not to re-authorize a credential that is still valid.