π Day 222 β Moving a Message Safety Filter to the Final Delivery Step
π Topic
During an in-progress Hermes WhatsApp repair, I found that filtering assistant responses was not enough. Direct handlers and maintenance notices could bypass that path and reach the delivery adapter. The defensive control had to sit at the final outbound boundary.
π― Goal
Prevent non-owner WhatsApp contacts from receiving internal model or capability notices, while preserving owner-only diagnostics and avoiding claims about images that the system did not inspect.
π What I Did
This is an active repair, so I am deliberately not calling its end-to-end validation complete yet. The implementation work so far has:
- identified a path where direct handlers could produce a reply without passing through the usual agent-response sanitizer;
- added an outbound-boundary check at the adapter path, where every final WhatsApp response is about to be delivered;
- used bridge-authenticated owner metadata to preserve narrowly scoped owner diagnostics while suppressing internal image/capability notices for non-owners;
- changed non-owner, non-vision image handling to an intentional no-reply instead of pretending to have inspected pixels or exposing plumbing prose;
- added regression coverage for that decision and for message pacing based on a wall-clock deadline;
- changed pacing to recheck short wall-clock intervals because a laptop suspension can distort a single long monotonic sleep.
The remaining boundary is important: the final full test rerun and live-safe validation were not yet recorded when this post was drafted.
π Key Cybersecurity Connections
Security filtering must guard the path an untrusted or unusual input can actually take. A policy in the ordinary response builder is useful, but it is not the final control if plugins, direct handlers, status callbacks, or retry paths can bypass it. This is defense in depth with a clear choke point: sanitize content again where it crosses from internal machinery into a human-facing channel.
The no-reply choice is also a privacy control. If the model has no visual capability, an automatic reaction or explanation can disclose internal implementation details or imply perception it never had. Suppressing the unsafe response is better than creating plausible but misleading text.
π Investigation Questions
- Which code paths can emit user-facing content without using the normal response builder?
- Is the final delivery adapter enforcing the same privacy policy as upstream logic?
- How is owner-only diagnostic authority authenticated at the boundary?
- Does a system ever claim it saw media it did not process?
- Can sleep, retry, or suspend behavior make reply pacing violate the intended wall-clock policy?
π¨ Detection Opportunities
Look for:
- direct handlers that skip a shared sanitizer;
- internal model/provider/capability notices reaching a personal chat surface;
- a metadata flag that is trusted without authentication;
- replies reacting to unavailable image content;
-
long timing delays that behave differently across sleep and resume.
event=outbound_policy_bypass source=direct_handler_or_status_callback action=apply_policy_at_final_delivery_boundary
π§ MITRE ATT&CK Techniques
No direct ATT&CK mapping claimed. This is privacy-preserving message delivery and boundary design, not a specific adversary-technique analysis.
πΊ Visual Investigation Diagram
Inbound event
β
Agent path OR direct handler/status callback
β
Some paths bypass ordinary response shaping
β
Final WhatsApp delivery boundary
β
Owner-authenticated diagnostic exception / non-owner privacy suppression
β
Safe textβor an intentional no-reply
β Challenges
The trap was assuming that a correct filter in the main path covered all replies. It did not. The real system had more than one writer, so the final delivery point needed to protect against all of them.
π What I Learned
I learned to map every user-visible output route before calling a content policy complete. A control is only effective on the surface that actually delivers the bytes.
β‘ Next Steps
- Run and record the final focused test set.
- Validate the boundary with a safe owner-controlled harness, not an unaware contact.
- Audit other direct output paths for the same bypass pattern.
π§ Reflection
The most valuable result was not a new regular expression. It was finding the correct enforcement location: the one place internal text becomes a message another person can read.
π§© Lessons Learned
What worked
Moving the policy to the final delivery boundary and making the non-owner non-vision path intentionally quiet.
What remains unverified
Final test and live-safe validation evidence for this in-progress repair.
Fix / takeaway
Put human-facing privacy controls at the final egress boundary, then prove every unusual output path reaches it.
π Skill Progression Context
This deepens my understanding of secure messaging design: output filtering, authenticated exceptions, truthful capability behavior, and safe validation matter together on personal communication surfaces.
π TL;DR
Filtering the main reply path was not enough because direct handlers could bypass it. The safer control belongs at the final WhatsApp delivery boundaryβand this repair stays labeled in progress until its final validation is recorded.
