🔄 Topic

Two real incidents — “Buonanotte zia!” sent at 11:24 in the morning, and “Buongiorno zia!” sent at 13:41 in the afternoon — traced back to a stale cached fact baked into the persona’s config by an hourly cron job. In the same pass I fixed a separate, unrelated resilience gap: three real turns had been silently dropped by a restart issued while work was still in flight.


🎯 Goal

Stop the persona from asserting a time-of-day fact that can be hours stale, and stop restarts from silently discarding active work.


🛠 What I Did

I traced two live, user-visible incidents to their actual root causes and fixed both at the source rather than patching the symptom.

Main areas covered:

  • traced “Buonanotte zia!” at 11:24 and “Buongiorno zia!” at 13:41 to a single cause: the persona carried a baked-in clock fact (“Oggi è …, sono le HH:MM.”) written into config.yaml by an hourly cron job
  • found the actual staleness window was worse than “up to an hour”: the change-detection logic (comparable()) deliberately ignores clock-only changes so they never force a restart on their own — meaning the clock fact could sit unrefreshed for far longer than the nominal hourly cadence, until something else happened to trigger an update
  • moved the clock fact out of the cron-refreshed config entirely and into a live, per-request block generated at the moment of actual use, in the gateway itself — so “what time is it” is answered live, not read from a cache that has no freshness guarantee tied to its own claim
  • left the cron script owning only the calendar facts it was always meant to own, since those genuinely only need hourly freshness and don’t carry the same failure mode
  • separately, investigated three real turns silently dropped on an earlier date by restarts issued while active_agents was still greater than zero — work was in flight, and the restart discarded it without any visible error
  • found the fix already existed in one place (this script’s own restart path already waited for the system to go idle before restarting) but wasn’t reusable
  • built hermes-safe-restart, generalizing that same proven pattern — wait for idle, then restart, judge success by an observed process-id change rather than trusting the exit code — into a shared script anyone can reach for instead of the raw, drain-timeout-dependent restart command

🔗 Key Cybersecurity Connections

A cached fact with no freshness guarantee is a data-integrity bug with a very literal blast radius here: the system asserted something false (the time of day) directly to a real person, twice, in a way that was immediately, obviously wrong to the recipient. The deliberate design choice to ignore clock-only changes in the restart-trigger logic is a good instinct in isolation — you don’t want to restart a whole persona just because a minute ticked over — but it quietly created an unbounded staleness window for a fact with no other refresh path. The fix is the right shape: don’t cache what needs to be live: compute it at the point of use instead of trusting a snapshot to still be true.

The dropped-turns bug is a race condition between “restart” and “work in progress,” and judging restart success by an observed process change rather than a trusted exit code is the more defensible verification: an exit code tells you the command didn’t error, not that the intended effect actually happened — the same “verify the real-world effect, not the absence of an error” discipline that’s shown up repeatedly this month.


🔍 Investigation Questions

  • Does any fact the system asserts to a user come from a cache with no freshness guarantee matching how often it’s claimed to be accurate?
  • Does change-detection logic that deliberately ignores certain fields (like clock-only changes) also account for what else was supposed to keep those fields fresh?
  • Is a proven safety pattern (wait for idle before restarting) actually reusable, or does it only protect the one script it was originally written into?
  • Does restart success get judged by an observed effect (a process actually changing) or by an exit code that only proves the command didn’t error?
  • How many turns, historically, were silently dropped by this exact race condition before it was found?

🚨 Detection Opportunities

Checks for cached-fact staleness and restart-safety:

  • a user-facing fact sourced from a periodically-refreshed cache with no freshness guarantee tied to how it’s presented
  • change-detection logic that excludes a field from triggering an update, without confirming something else refreshes that field independently
  • a restart issued without checking for in-flight work, risking silently dropped turns
  • restart success judged by exit code alone rather than an observed state change
  • a proven safety pattern (like wait-for-idle) implemented once and never made reusable elsewhere

Example:

project=hermes-whatsapp-persona
signal=cached_fact_asserted_stale_to_real_user
risk_area=unbounded_cache_staleness_window
triage=move_time_sensitive_facts_to_live_per_request_computation

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is a data-integrity (cache staleness) and reliability (restart race condition) issue, not an adversary technique.


🗺 Visual Investigation Diagram

Persona says "Buonanotte" at 11:24am
    ↓
Clock fact baked into config.yaml, hourly cron refresh
    ↓
Change-detection ignores clock-only changes — no forced refresh
    ↓
Staleness window unbounded in practice
    ↓
Fix: compute clock fact live, per request, in the gateway
    ↓
Separately: three turns dropped by restart during active work
    ↓
hermes-safe-restart: wait for idle, restart, verify by observed pid change

⚠ Challenges

The clock bug was deceptively simple-looking — “just refresh it more often” would have been the easy, wrong fix, since the real cause was that clock-only changes were deliberately excluded from triggering any refresh at all. The right fix meant recognizing that a live, time-sensitive fact has no business living in a periodically-refreshed cache in the first place, no matter how often that cache runs.


📚 What I Learned

I learned that “refresh more often” is a bandage over a caching bug, not a fix — a fact that must always be current shouldn’t be cached at all, it should be computed at the point of use. I also learned that a good, proven safety pattern is only actually protective everywhere it’s needed once it’s extracted into something reusable, rather than living correctly in exactly one script and nowhere else.


➡ Next Steps

  • Audit other config-baked facts for the same clock-only-exclusion staleness risk
  • Roll hermes-safe-restart out as the default restart path everywhere the raw command is still used
  • Add a regression test asserting the clock block is computed live, not read from a cached config value
  • Review historical logs for other instances of turns silently dropped by the same restart race condition

🧠 Reflection

Two people getting a “good morning” in the afternoon is a small, almost funny-sounding bug — but it’s a direct, visible instance of the system stating something false with full confidence, which is exactly the failure shape worth taking seriously regardless of how low the stakes look on the surface.


🧩 Lessons Learned

What worked

Moving the time-sensitive fact to live, per-request computation instead of a periodically-refreshed cache, and generalizing the wait-for-idle restart pattern into a reusable script.

What broke

The persona sent factually wrong time-of-day greetings twice due to unbounded cache staleness, and three real turns were silently dropped by a restart issued mid-work.

Why it broke

A field was deliberately excluded from triggering a refresh, with no other mechanism keeping it current, and a restart pattern that correctly waited for idle in one script wasn’t reusable anywhere else.

Fix / takeaway

Don’t cache what must always be current — compute it live at the point of use — and extract proven safety patterns into reusable tools rather than leaving them correct in only one place.


📈 Skill Progression Context

This supports my cybersecurity progression because cache-staleness analysis, race-condition identification around restarts, and verifying real-world effect rather than trusting exit codes are core data-integrity and reliability engineering skills that directly support security work.


😄 TL;DR

The bot wished someone good night at 11:24am because of a stale cached clock fact — fixed by making time live instead of cached, and while I was in there, made sure restarts stop silently dropping work in progress.