📅 Day 205 — Fixing a Bot That Said Good Night at 11:24 in the Morning (Stale Cached Time)
🔄 Topic
Two real incidents — “Buonanotte zia!” sent at 11:24 in the morning, and “Buongiorno zia!” sent at 13:41 in the afternoon — traced back to a stale cached fact baked into the persona’s config by an hourly cron job. In the same pass I fixed a separate, unrelated resilience gap: three real turns had been silently dropped by a restart issued while work was still in flight.
🎯 Goal
Stop the persona from asserting a time-of-day fact that can be hours stale, and stop restarts from silently discarding active work.
🛠 What I Did
I traced two live, user-visible incidents to their actual root causes and fixed both at the source rather than patching the symptom.
Main areas covered:
- traced “Buonanotte zia!” at 11:24 and “Buongiorno zia!” at 13:41 to a single cause: the persona carried a baked-in clock fact (“Oggi è …, sono le HH:MM.”) written into
config.yamlby an hourly cron job - found the actual staleness window was worse than “up to an hour”: the change-detection logic (
comparable()) deliberately ignores clock-only changes so they never force a restart on their own — meaning the clock fact could sit unrefreshed for far longer than the nominal hourly cadence, until something else happened to trigger an update - moved the clock fact out of the cron-refreshed config entirely and into a live, per-request block generated at the moment of actual use, in the gateway itself — so “what time is it” is answered live, not read from a cache that has no freshness guarantee tied to its own claim
- left the cron script owning only the calendar facts it was always meant to own, since those genuinely only need hourly freshness and don’t carry the same failure mode
- separately, investigated three real turns silently dropped on an earlier date by restarts issued while
active_agentswas still greater than zero — work was in flight, and the restart discarded it without any visible error - found the fix already existed in one place (this script’s own restart path already waited for the system to go idle before restarting) but wasn’t reusable
- built
hermes-safe-restart, generalizing that same proven pattern — wait for idle, then restart, judge success by an observed process-id change rather than trusting the exit code — into a shared script anyone can reach for instead of the raw, drain-timeout-dependent restart command
🔗 Key Cybersecurity Connections
A cached fact with no freshness guarantee is a data-integrity bug with a very literal blast radius here: the system asserted something false (the time of day) directly to a real person, twice, in a way that was immediately, obviously wrong to the recipient. The deliberate design choice to ignore clock-only changes in the restart-trigger logic is a good instinct in isolation — you don’t want to restart a whole persona just because a minute ticked over — but it quietly created an unbounded staleness window for a fact with no other refresh path. The fix is the right shape: don’t cache what needs to be live: compute it at the point of use instead of trusting a snapshot to still be true.
The dropped-turns bug is a race condition between “restart” and “work in progress,” and judging restart success by an observed process change rather than a trusted exit code is the more defensible verification: an exit code tells you the command didn’t error, not that the intended effect actually happened — the same “verify the real-world effect, not the absence of an error” discipline that’s shown up repeatedly this month.
🔍 Investigation Questions
- Does any fact the system asserts to a user come from a cache with no freshness guarantee matching how often it’s claimed to be accurate?
- Does change-detection logic that deliberately ignores certain fields (like clock-only changes) also account for what else was supposed to keep those fields fresh?
- Is a proven safety pattern (wait for idle before restarting) actually reusable, or does it only protect the one script it was originally written into?
- Does restart success get judged by an observed effect (a process actually changing) or by an exit code that only proves the command didn’t error?
- How many turns, historically, were silently dropped by this exact race condition before it was found?
🚨 Detection Opportunities
Checks for cached-fact staleness and restart-safety:
- a user-facing fact sourced from a periodically-refreshed cache with no freshness guarantee tied to how it’s presented
- change-detection logic that excludes a field from triggering an update, without confirming something else refreshes that field independently
- a restart issued without checking for in-flight work, risking silently dropped turns
- restart success judged by exit code alone rather than an observed state change
- a proven safety pattern (like wait-for-idle) implemented once and never made reusable elsewhere
Example:
project=hermes-whatsapp-persona
signal=cached_fact_asserted_stale_to_real_user
risk_area=unbounded_cache_staleness_window
triage=move_time_sensitive_facts_to_live_per_request_computation
🧭 MITRE ATT&CK Techniques
No direct mapping claimed. This is a data-integrity (cache staleness) and reliability (restart race condition) issue, not an adversary technique.
🗺 Visual Investigation Diagram
Persona says "Buonanotte" at 11:24am
↓
Clock fact baked into config.yaml, hourly cron refresh
↓
Change-detection ignores clock-only changes — no forced refresh
↓
Staleness window unbounded in practice
↓
Fix: compute clock fact live, per request, in the gateway
↓
Separately: three turns dropped by restart during active work
↓
hermes-safe-restart: wait for idle, restart, verify by observed pid change
⚠ Challenges
The clock bug was deceptively simple-looking — “just refresh it more often” would have been the easy, wrong fix, since the real cause was that clock-only changes were deliberately excluded from triggering any refresh at all. The right fix meant recognizing that a live, time-sensitive fact has no business living in a periodically-refreshed cache in the first place, no matter how often that cache runs.
📚 What I Learned
I learned that “refresh more often” is a bandage over a caching bug, not a fix — a fact that must always be current shouldn’t be cached at all, it should be computed at the point of use. I also learned that a good, proven safety pattern is only actually protective everywhere it’s needed once it’s extracted into something reusable, rather than living correctly in exactly one script and nowhere else.
➡ Next Steps
- Audit other config-baked facts for the same clock-only-exclusion staleness risk
- Roll
hermes-safe-restartout as the default restart path everywhere the raw command is still used - Add a regression test asserting the clock block is computed live, not read from a cached config value
- Review historical logs for other instances of turns silently dropped by the same restart race condition
🧠 Reflection
Two people getting a “good morning” in the afternoon is a small, almost funny-sounding bug — but it’s a direct, visible instance of the system stating something false with full confidence, which is exactly the failure shape worth taking seriously regardless of how low the stakes look on the surface.
🧩 Lessons Learned
What worked
Moving the time-sensitive fact to live, per-request computation instead of a periodically-refreshed cache, and generalizing the wait-for-idle restart pattern into a reusable script.
What broke
The persona sent factually wrong time-of-day greetings twice due to unbounded cache staleness, and three real turns were silently dropped by a restart issued mid-work.
Why it broke
A field was deliberately excluded from triggering a refresh, with no other mechanism keeping it current, and a restart pattern that correctly waited for idle in one script wasn’t reusable anywhere else.
Fix / takeaway
Don’t cache what must always be current — compute it live at the point of use — and extract proven safety patterns into reusable tools rather than leaving them correct in only one place.
📈 Skill Progression Context
This supports my cybersecurity progression because cache-staleness analysis, race-condition identification around restarts, and verifying real-world effect rather than trusting exit codes are core data-integrity and reliability engineering skills that directly support security work.
😄 TL;DR
The bot wished someone good night at 11:24am because of a stale cached clock fact — fixed by making time live instead of cached, and while I was in there, made sure restarts stop silently dropping work in progress.
