๐ Day 236 โ Testing a Messaging Bot's Final Message Instead of the Model's First Output
๐ Topic
I spent time testing a local messaging persona with tools attached. The most important discovery was about the measuring instrument: scoring the first model completion was not the same as scoring the message that a person would receive.
๐ฏ Goal
Build a test that follows the real path far enough to answer a security question: after allowed tools, retries, filtering, and shaping have run, what actually reaches the delivery boundary?
๐ What I Did
The first version of the probe treated the first completion as the final answer. That was correct when the surface had no usable tools, but it became misleading after the tool loop was enabled. A tool call could execute successfully and produce a later, human-readable answer. Counting the first step as failure made the model look worse than the delivered behavior.
I changed the test to carry permitted tool calls through bounded rounds and to score the resulting text. Unknown tool calls still fail immediately, and a loop that never returns to text remains a real failure because it would produce silence.
I also made the test use the same kinds of transient context that production supplies: the live clock, the generated agenda block, and the per-chat steering data. Without that context, a supposedly positive agenda assertion could never fire. The test was asking a narrower question than the production system.
The comparison was deliberately local and owner-controlled. No unaware person received a test message. Across repeated runs, the raw failure totals moved from 30 to 13 and then 16. I cannot call that a stable trend from three runs. The more useful result was at the delivery boundary: after classification, trimming, and shaping, far fewer failures ended as silence, and the delivered samples no longer included the self-disclosure that the old filter had missed.
๐ Key Cybersecurity Connections
Security tests can fail in two directions:
- a real defect can be hidden by a harness that checks the wrong stage
- a correct response can be labelled bad because the harness stops too early
This is the same distinction as testing an authentication system by checking a configuration file instead of asking the running service what decision it makes. The test must exercise the boundary that matters.
๐ Investigation Questions
- Does the test send the same essential context as production?
- Does it follow an allowed tool call to the final text, within a defined limit?
- Does it include both positive and negative controls?
- Does it score the final transformed output rather than an intermediate object?
- Can a known-bad fixture make the harness fail?
๐จ Detection Opportunities
Useful signals include a tool call followed by no final text, a response that changes after filtering, a harness that never exercises a positive control, and a failure count that changes merely because the probe truncates the response.
Example investigation record:
test_stage=delivery_boundary
input=allowed_tool_request
tool_rounds=2
final_text=present
filter_verdict=delivered
shaping_applied=false
๐งญ MITRE ATT&CK Techniques
No direct ATT&CK mapping is claimed. This post is about validating a defensive control and its measurement boundary, not evidence that an adversary technique occurred.
๐บ Visual Investigation Diagram
Scenario fixture
โ
Production-shaped prompt and context
โ
Bounded model/tool loop
โ
Classify outgoing text
โ
Shape or suppress where policy requires
โ
Score the message a person would receive
โ Challenges
The difficult part was accepting that a green test suite can still answer the wrong question. The probe had valid assertions, but one of them was placed before the systemโs important transition. A failure count looked precise while its meaning had changed underneath it.
๐ What I Learned
โThe model repliedโ is not a sufficient security observation. I need to know which completion, after which tools, under which context, and after which egress controls. Intermediate results are useful diagnostics, but they are not delivery evidence.
โก Next Steps
- Keep positive and negative controls in every output-safety suite
- Re-score measurements after a capability change
- Record whether a result describes generation, filtering, shaping, or delivery
- Repeat the same scenario set before interpreting a number as progress
๐ง Reflection
The best improvement was not a new prompt rule. It was making the test follow the system far enough to expose what a person would actually see. A security control is only as trustworthy as the path the test observes.
๐งฉ Lessons Learned
What worked
Following allowed tool calls through a bounded loop and checking the final transformed text.
What could go wrong
Counting an intermediate completion as the outcome and silently disabling a positive assertion by omitting required context.
Takeaway
Measure the security boundary that matters, not the easiest object to inspect.
๐ Skill Progression Context
This moves my cybersecurity learning from writing rules to validating controls. It is the same evidence discipline I am learning in vulnerability assessment: name the question, define the scope, and avoid claiming that a narrower test proves a broader result.
๐ TL;DR
A modelโs first completion is not necessarily the message a person receives. Security tests must follow the real path to the delivery boundary.
