π Day 238 β Extending an Outgoing-Message Guard to Check the Message It Replies To
π Topic
I extended an outgoing-message guard after discovering that some failures cannot be detected from the reply alone. Whether a message is safe can depend on the inbound message it answers, the language it used, and whether a closing sentence changes its meaning.
π― Goal
Make the last delivery boundary inspect the final content and the relevant relationship to the input, then shape recoverable problems without weakening rules that must remain fail-closed.
π What I Did
The original filter classified outgoing text as if it were independent of the message that triggered it. That missed two relational signals: an answer in English to an Italian message and an apologetic, service-style response to an insult. The guard now receives the triggering message from the bridgeβs own recent history and can evaluate the pair.
I also separated problems that can be safely shaped from problems that must be suppressed. Lists, markdown-style formatting, self-quoting, and some long trailing service offers can be transformed or trimmed and then classified again. Leaked reasoning, identity admissions, tracebacks, and stray script tokens are not treated as cosmetic. If they remain, the complete message is rejected.
The important word is again. A shaped message is not automatically safe because the first classifier liked it. The shortened result must pass through the relevant checks a second time, and there must be a minimum amount of useful text left.
The test path remained local and owner-controlled. No real contact was used as a harness. The final measurement showed that shaping changed what reached the boundary: some register failures arrived as short, human-sounding text instead of either an essay or silence. The model still had occasional weaknesses, so I recorded the limitation rather than calling the system solved.
π Key Cybersecurity Connections
This is an egress-control pattern. A source component may generate something that is technically valid text but unsafe for the destination. The final adapter is where policy must hold because every upstream component is allowed to make a mistake.
π Investigation Questions
- Does the guard know which input the output answers?
- Is the final text checked after trimming, splitting, or formatting?
- Which findings are safe to transform and which require suppression?
- Can an apparently harmless closing offer change the register of the whole reply?
- Does a fallback preserve the same privacy and authorization rules?
π¨ Detection Opportunities
Look for language mismatches, assistant-register markers, leaked internal reasoning, unexpected markdown, quotation wrappers around an entire reply, and output that changes materially during shaping. These signals should be logged with classification reason and final disposition without recording private message bodies.
Example:
relation=reply_to_inbound
input_language=it
output_language=en
action=review_or_suppress
body_logging=disabled
π§ MITRE ATT&CK Techniques
No direct ATT&CK mapping is claimed. This is a defensive output-integrity and privacy boundary. ATT&CK techniques would only be selected if later telemetry showed adversary behavior.
πΊ Visual Investigation Diagram
Inbound message metadata
β
Model output
β
Classify content + relationship
β
Trim / split / shape only allowed classes
β
Re-classify the final text
β
Deliver or suppress
β Challenges
It is tempting to fix every bad reply with another prompt sentence. That is not enough at the last boundary. Some failures are about content; others are about the relationship between two messages. A stateless regular expression cannot answer a relational question.
π What I Learned
Security filters should be attached to the surface where the failure becomes harmful. They should also distinguish βcan be made safeβ from βmust not pass.β A guard that treats a credential leak like a formatting problem is not a guard.
β‘ Next Steps
- Keep body logging out of metadata-only detection records
- Add paired fixtures: true relational triggers and near-neighbour non-triggers
- Exercise every shaping step with a second classification pass
- Recheck the actual delivery pipeline after every policy change
π§ Reflection
The strongest control was not the longest list of forbidden phrases. It was the combination of context, final-boundary placement, explicit fail-closed classes, and a second check after transformation.
π§© Lessons Learned
What worked
Passing minimal inbound context into the egress decision and rechecking shaped output.
What could go wrong
Treating a classification result for an intermediate string as proof about the final message.
Takeaway
Validate the exact artifact that crosses the boundary.
π Skill Progression Context
This builds on my earlier work with evidence-gated delivery and connects it to secure software design: context, authorization, transformation, and final verification must compose without creating a bypass.
π TL;DR
The last message check needs enough context to understand the reply, and it must inspect the final shaped text before delivery.
