🔄 Topic

Rolling WhatsApp out as a real channel for Hermes surfaced two unrelated ways a new surface can end up trusting more than it should: a default toolset that handed a chat-facing persona terminal and code execution, and a greeting feature that couldn’t tell “the group is active” from “Juri himself said something.”


🎯 Goal

Give the new WhatsApp channel exactly the tools and exactly the identity checks it needs — no more — and fix both gaps found while actually watching it run against real family traffic.


🛠 What I Did

I audited a new channel from two different angles: what it’s allowed to do, and who it correctly recognizes as whom.

Main areas covered:

  • found the default platform_toolsets bundle for WhatsApp exposed terminal, computer_use, execute_code, and file write to anyone in the loop — capabilities that made sense for a CLI-facing agent but not for a persona replying to family members over chat
  • locked platform_toolsets.whatsapp down to [vision] only, matching the actual, narrow job the channel needs to do
  • fixed a separate timing bug in the same pass: a randomized morning-greeting send window had a gate bug that let a “Buongiorno” fire at 14:29 in the afternoon, and would have fired again via cron on a late-wake day
  • shipped a feature to skip the morning greeting in a chat that’s already organically “awake” — if someone had already messaged and been replied to, a scripted good-morning on top would be redundant
  • found that logic was wrong for groups: on a real morning, my parents both posted “Buongiorno” in the family group, the greeting was suppressed because the chat looked “active,” and I — who had said nothing — simply never appeared to reply. Twelve people saw both parents greet the group and no reply from me.
  • traced the root cause: the skip logic was written for DMs, where any message really is an exchange with me — she writes, I answer. A group is not that; other people talking to each other is not me participating, and suppressing on their activity produces exactly the silence the feature was supposed to prevent
  • fixed it so groups require my own outbound message to suppress the greeting, while DMs keep counting either direction — verified against the day’s real activity data across a group with only inbound messages, a group with my own outbound message, and a DM with only the other person’s message
  • also added greeted-chat tracking so the agent stops repeating the same greeting, and fixed a bug where “one person” was wrongly treated as “one chat id” when a contact has multiple entry points

🔗 Key Cybersecurity Connections

The toolset bundle is a classic least-privilege gap: a capability set built for one trust context (an agent with a human directly steering it) got inherited by a different, lower-trust context (a persona replying autonomously in a family chat) without being re-scoped. The fix isn’t clever — it’s just actually asking “what does this surface need” instead of reusing a default.

The greeting bug is a subtler identity problem: the system conflated “activity in a channel” with “activity by the specific person the feature cares about.” That’s the same shape of mistake as conflating “a request came from an authenticated session” with “a request came from the specific user that session claims to be” — proximity or context isn’t identity, and a control built on the wrong proxy for identity fails exactly when it matters, silently, in front of an audience (in this case, literally twelve people).


🔍 Investigation Questions

  • Does every channel get its own scoped toolset, or does it inherit a default built for a different trust context?
  • Does an “activity” or “presence” signal actually mean the specific person the feature cares about was active — or just that the channel was active in general?
  • Would a timing or gating bug (like a send window) actually be caught by normal testing, or only by watching real-world execution?
  • Is a chat-identity mapping (one person → one chat id) verified against reality, or assumed to be 1:1?
  • When a feature fails silently and visibly to an audience, is the root cause traced to its actual design assumption, or just patched at the symptom?

🚨 Detection Opportunities

Checks for a new communication channel’s trust boundaries:

  • a new channel inheriting a default toolset built for a different, higher-trust context
  • a presence/activity signal used as a proxy for a specific individual’s identity without verification
  • a send-window or timing gate untested against real day-boundary and late-wake conditions
  • a chat-identity mapping assumed 1:1 without checking for multiple entry points per contact
  • a feature failing silently in a way visible to end users before it’s caught internally

Example:

project=hermes-whatsapp-channel
signal=default_toolset_inherited_without_rescoping
risk_area=excess_privilege_on_new_communication_surface
triage=enumerate_actual_required_capabilities_scope_explicitly

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is least-privilege scoping and identity-vs-proxy verification for a new communication channel, not an adversary technique.


🗺 Visual Investigation Diagram

New WhatsApp channel rolled out
    ↓
Default toolset inherited: terminal, code exec, file write exposed
    ↓
Rescoped to [vision] only — least privilege for the actual job
    ↓
Separately: greeting suppressed on "chat activity"
    ↓
Parents post, Juri silent, greeting wrongly suppressed
    ↓
Root cause: activity ≠ Juri's own activity, esp. in groups
    ↓
Groups now require his own outbound message to suppress

⚠ Challenges

The greeting bug was uncomfortable specifically because of how it failed: not with an error, not silently in a log nobody reads, but visibly, in front of family, on a real morning. That’s a good reminder that “silent failure” doesn’t mean “unnoticed failure” — it means the system itself doesn’t notice, even when people very much do.


📚 What I Learned

I learned that a default configuration inherited by a new, lower-trust surface is a security decision even when nobody explicitly made one — the absence of explicit scoping is itself the vulnerability. I also learned that “activity” is a dangerously convenient but often wrong proxy for “this specific person’s activity,” especially the moment a feature moves from one-on-one contexts into groups.


➡ Next Steps

  • Audit other channels for inherited-but-unscoped toolsets, not just WhatsApp
  • Add a regression test asserting groups only suppress on the account’s own outbound message
  • Review other “activity”-based signals in the stack for the same person-vs-channel conflation
  • Re-verify the send-window gate against more day-boundary and late-wake scenarios

🧠 Reflection

Two unrelated bugs, found in the same rollout, taught the same underlying lesson from different angles: trust has to be explicitly scoped to the actual context, not inherited from wherever it happened to come from — whether that’s a tool bundle or a “was this chat active” signal standing in for “was he active.”


🧩 Lessons Learned

What worked

Explicitly auditing both the toolset scope and the identity logic for a new channel, rather than assuming either generalized correctly from an existing pattern.

What broke

A WhatsApp persona had terminal, code-exec, and file-write access it never needed, a send-window gate fired hours late, and a group greeting suppressed itself on other people’s activity, leaving Juri visibly silent to his own family.

Why it broke

Both bugs came from reusing an assumption from a different, lower-stakes context — a shared default toolset, and a DM-shaped activity check applied to groups — without re-verifying it fit the new one.

Fix / takeaway

Scope every new channel’s capabilities explicitly rather than inheriting defaults, and treat “activity” as a proxy for identity that needs verifying, especially the moment a feature moves from one-to-one into a group setting.


📈 Skill Progression Context

This supports my cybersecurity progression because least-privilege scoping of new capabilities and distinguishing identity from proximate signals (like channel activity) are foundational access-control and authentication concepts, exercised here on a live system with a real, visible failure mode.


😄 TL;DR

The new WhatsApp channel had too many tools and too little judgment about who was actually talking — locked the toolset down to what it needs, and taught the greeting logic that “the group is active” isn’t the same as “I said something.”