๐Ÿ”„ Topic

I spent time testing a local messaging persona with tools attached. The most important discovery was about the measuring instrument: scoring the first model completion was not the same as scoring the message that a person would receive.

๐ŸŽฏ Goal

Build a test that follows the real path far enough to answer a security question: after allowed tools, retries, filtering, and shaping have run, what actually reaches the delivery boundary?

๐Ÿ›  What I Did

The first version of the probe treated the first completion as the final answer. That was correct when the surface had no usable tools, but it became misleading after the tool loop was enabled. A tool call could execute successfully and produce a later, human-readable answer. Counting the first step as failure made the model look worse than the delivered behavior.

I changed the test to carry permitted tool calls through bounded rounds and to score the resulting text. Unknown tool calls still fail immediately, and a loop that never returns to text remains a real failure because it would produce silence.

I also made the test use the same kinds of transient context that production supplies: the live clock, the generated agenda block, and the per-chat steering data. Without that context, a supposedly positive agenda assertion could never fire. The test was asking a narrower question than the production system.

The comparison was deliberately local and owner-controlled. No unaware person received a test message. Across repeated runs, the raw failure totals moved from 30 to 13 and then 16. I cannot call that a stable trend from three runs. The more useful result was at the delivery boundary: after classification, trimming, and shaping, far fewer failures ended as silence, and the delivered samples no longer included the self-disclosure that the old filter had missed.

๐Ÿ”— Key Cybersecurity Connections

Security tests can fail in two directions:

  • a real defect can be hidden by a harness that checks the wrong stage
  • a correct response can be labelled bad because the harness stops too early

This is the same distinction as testing an authentication system by checking a configuration file instead of asking the running service what decision it makes. The test must exercise the boundary that matters.

๐Ÿ” Investigation Questions

  • Does the test send the same essential context as production?
  • Does it follow an allowed tool call to the final text, within a defined limit?
  • Does it include both positive and negative controls?
  • Does it score the final transformed output rather than an intermediate object?
  • Can a known-bad fixture make the harness fail?

๐Ÿšจ Detection Opportunities

Useful signals include a tool call followed by no final text, a response that changes after filtering, a harness that never exercises a positive control, and a failure count that changes merely because the probe truncates the response.

Example investigation record:

test_stage=delivery_boundary
input=allowed_tool_request
tool_rounds=2
final_text=present
filter_verdict=delivered
shaping_applied=false

๐Ÿงญ MITRE ATT&CK Techniques

No direct ATT&CK mapping is claimed. This post is about validating a defensive control and its measurement boundary, not evidence that an adversary technique occurred.

๐Ÿ—บ Visual Investigation Diagram

Scenario fixture
    โ†“
Production-shaped prompt and context
    โ†“
Bounded model/tool loop
    โ†“
Classify outgoing text
    โ†“
Shape or suppress where policy requires
    โ†“
Score the message a person would receive

โš  Challenges

The difficult part was accepting that a green test suite can still answer the wrong question. The probe had valid assertions, but one of them was placed before the systemโ€™s important transition. A failure count looked precise while its meaning had changed underneath it.

๐Ÿ“š What I Learned

โ€œThe model repliedโ€ is not a sufficient security observation. I need to know which completion, after which tools, under which context, and after which egress controls. Intermediate results are useful diagnostics, but they are not delivery evidence.

โžก Next Steps

  • Keep positive and negative controls in every output-safety suite
  • Re-score measurements after a capability change
  • Record whether a result describes generation, filtering, shaping, or delivery
  • Repeat the same scenario set before interpreting a number as progress

๐Ÿง  Reflection

The best improvement was not a new prompt rule. It was making the test follow the system far enough to expose what a person would actually see. A security control is only as trustworthy as the path the test observes.

๐Ÿงฉ Lessons Learned

What worked

Following allowed tool calls through a bounded loop and checking the final transformed text.

What could go wrong

Counting an intermediate completion as the outcome and silently disabling a positive assertion by omitting required context.

Takeaway

Measure the security boundary that matters, not the easiest object to inspect.

๐Ÿ“ˆ Skill Progression Context

This moves my cybersecurity learning from writing rules to validating controls. It is the same evidence discipline I am learning in vulnerability assessment: name the question, define the scope, and avoid claiming that a narrower test proves a broader result.

๐Ÿ˜„ TL;DR

A modelโ€™s first completion is not necessarily the message a person receives. Security tests must follow the real path to the delivery boundary.