📅 Day 203 — A Config Cap That Deadlocked the Gateway, and a Prefix Match That Nearly Revoked the Wrong User
🔄 Topic
One day carried three separate lessons: a compression-threshold cap meant to fix a real problem instead deadlocked the whole gateway and had to be reverted, an access-revocation for a user I’d asked to stop turned into a near-miss because of how phone numbers get matched, and a “claim detector” got built to catch the assistant asserting things that weren’t actually true.
🎯 Goal
Fix a real latency problem without breaking the system worse than the problem itself, revoke access for the right person and only that person, and start automatically catching a category of bug — false claims in the assistant’s own output — before it ships instead of after.
🛠 What I Did
I shipped a fix, watched it fail worse than the original bug, reverted it, and used the same session to close two unrelated gaps.
Main areas covered:
- capped
agent.compression.threshold_tokensat 16000, aiming to fix a real problem: the ratio-based trigger sat at 48,742–64,000 tokens, well above a 38,658-token turn that had taken 15.6 minutes to resolve — compression had never fired once, all day, for exactly the turns that needed it most - the cap immediately deadlocked the gateway:
threshold_tokens: 16000was unreachable in practice becauseprotect_last_n: 20protects a tail that’s itself roughly 32K tokens of tool output — compression fired, freed only 117 tokens, logged insufficient progress, and fired again roughly every five minutes; a real turn hung for 16 minutes and blocked the deferred autoreply worker, which waits for the gateway to go idle - reverted the cap back to the ratio-based 48,742 trigger, and documented that the underlying 38,658-token problem still stands — the real fix has to start at
protect_last_nand tool-output size in history, not at the trigger threshold itself - separately revoked a user (referred to as “Aurora” in the handoff) from
WHATSAPP_ALLOWED_USERSafter being asked to stop contacting the bot — but resolved every entry to its canonical LID first, because a prefix match on phone number alone would have given a false picture: several Italian numbers share the same+39 379 1xxprefix, which is exactly how a revocation removes the wrong person - built a “claim detector” — a fabricated-action rule that flags when the assistant’s own output asserts something happened that the gateway didn’t actually observe. It detects but deliberately does not block, because the same sentence can be true or false depending on whether a tool actually ran, and only the gateway has that ground truth
- corrected two of my own earlier claims in the same handoff document: a second
llama-serverI’d flagged as stale was actually the tier proxy’s small model, and I’d been counting cache hits from a log line that doesn’t exist for cache hits — so the count itself was wrong
🔗 Key Cybersecurity Connections
Reverting a change because it made things worse, on the same day it shipped, is change management working correctly — the alternative (leaving a deadlocking fix in place because reverting feels like admitting failure) is how production incidents get prolonged. The revert also correctly separated symptom from cause: capping the trigger treated the number, not the actual reason compression wasn’t effective, which the write-up honestly leaves as still-open work.
The Aurora revocation is an identity-resolution problem with real consequences: matching on a raw phone-number prefix in a region where multiple people share a prefix range is the same failure class as matching on a truncated or non-unique identifier anywhere in a security control — resolve to a canonical, unique identity before any revocation, ban, or grant, or you risk acting on the wrong entity while believing you acted on the right one. The claim detector is a step toward the false-assurance pattern from earlier this month becoming something the system catches automatically instead of something I keep discovering after the fact — with an honest limitation stated up front: detection without a reliable way to block yet.
🔍 Investigation Questions
- Did the compression-threshold fix address the root cause, or just move where the same problem resurfaces?
- Was the fix reverted quickly once evidence showed it made things worse, or did it linger out of reluctance to undo shipped work?
- Does a revocation or ban resolve to a canonical unique identity first, or match on a prefix or partial identifier that could collide?
- Does the claim detector’s “detect but don’t block” limitation get tracked as open work, or forgotten once the detector exists?
- Were my own prior written claims (the stale-model note, the cache-hit count) re-verified rather than trusted because I’d written them down?
🚨 Detection Opportunities
Checks for change management and identity-resolution risk:
- a configuration change deployed without checking its interaction with an existing, related setting (here: threshold cap vs. protect_last_n)
- a fix left in production after evidence shows it caused a worse outage than the original problem
- a revocation, ban, or grant acting on a non-canonical or prefix-matched identifier in a region where collisions are possible
- a “detects but doesn’t block” control with no tracked follow-up to close the gap
- a previously written claim or measurement never re-verified before being relied on again
Example:
project=hermes-compression-tuning
signal=config_change_deadlocked_dependent_subsystem
risk_area=untested_interaction_between_related_settings
triage=revert_immediately_document_root_cause_still_open
🧭 MITRE ATT&CK Techniques
No direct mapping claimed. This is production change-management discipline and identity-resolution correctness for access revocation, not an adversary technique.
🗺 Visual Investigation Diagram
Compression trigger too high — real turns never compress
↓
Cap threshold_tokens at 16000
↓
Deadlock: protect_last_n tail alone exceeds new cap
↓
16-minute hang, blocks deferred autoreply worker
↓
Revert to ratio-based trigger, root cause left explicitly open
↓
Separately: revoke Aurora — resolve to canonical LID first
↓
Prefix-match would have revoked the wrong person
↓
Build claim detector: flags fabricated-action claims, doesn't yet block
⚠ Challenges
The hardest moment was recognizing the compression cap had made things worse within the same session it shipped, and reverting immediately rather than trying to tune around the deadlock to salvage the “fix.” The Aurora revocation was a quieter but sharper scare — a prefix match looked like it would work, and only checking canonical LIDs first caught that it would have hit an innocent person sharing a number prefix.
📚 What I Learned
I learned that a fix interacting badly with an existing, unrelated-looking setting (a trigger threshold and a token-protection window) is a real category of production risk, and that reverting fast on clear evidence is the correct response, not a failure. I also learned that identity resolution has to happen before any access-control action, every time, because “looks like the same number” is not the same claim as “is the same person.”
➡ Next Steps
- Fix the actual 38,658-token problem at its source — protect_last_n and tool-output size in history — rather than at the trigger
- Give the claim detector a path to actually block, not just flag, fabricated-action claims
- Audit other allowlist/denylist logic for prefix-based matching that should be canonical-identity-based instead
- Re-verify any other standing claims in handoff documents that haven’t been checked against current measurements
🧠 Reflection
Three unrelated near-misses in one session, and the common thread is the same one from earlier this month: verify before you trust — trust a threshold’s interaction with other settings, trust a prefix match as identity, trust your own prior written claims. None of them held up to a direct check, and checking is what caught each one before it became a real incident.
🧩 Lessons Learned
What worked
Reverting the compression cap immediately once it deadlocked the gateway, resolving to canonical LID before revoking access, and building a detector for a bug category instead of waiting to find the next instance by hand.
What broke
The compression-threshold cap deadlocked the gateway and hung a real turn for 16 minutes, a prefix match nearly revoked the wrong person, and two of my own earlier written claims turned out to be wrong on re-check.
Why it broke
The cap didn’t account for its interaction with protect_last_n, phone-number prefixes aren’t unique identities in this dataset, and written claims hadn’t been re-verified since first being logged.
Fix / takeaway
Revert fast on clear evidence rather than defending a shipped change, resolve to canonical identity before any access-control action, and treat your own prior claims as needing re-verification, not as settled fact.
📈 Skill Progression Context
This supports my cybersecurity progression because fast, evidence-based rollback discipline, canonical identity resolution before access-control actions, and building automated detectors for known failure patterns are all core to real-world security and reliability engineering.
😄 TL;DR
Capped a threshold to fix a real problem, watched it deadlock the gateway instead, reverted it same-day — then caught a prefix match that would have revoked the wrong person, and started building a detector for my own assistant’s false claims.
