Lab Objective

Build a safe local harness that tests the final output of a tool-using agent after bounded tool calls, filtering, and shaping. The goal is to avoid treating an intermediate model completion as the delivered outcome.

Lab Environment

  • Environment: local test harness and local output guard
  • Data: synthetic message fixtures and owner-controlled canary scenarios
  • Safety boundary: no real contact, external recipient, or private message body is used
  • Evidence: the final text, classification, shaping decision, and delivery disposition

Scenario

A model can return a tool call before it returns readable text. A filter can then trim or suppress the text. If the harness scores only the first response, it can report a false failure or miss a defect introduced later in the pipeline.

Step 1 - Define the Stages

Write down the stages before testing:

fixture → production-shaped context → bounded tool loop
        → classify → shape or suppress → classify again → disposition

The acceptance question is what remains at the final disposition, not merely what the model generated first.

Step 2 - Add Positive and Negative Controls

Use paired fixtures:

  • an allowed request that produces a normal final reply
  • an unknown tool call that must fail closed
  • a tool call followed by useful text
  • a reply containing a forbidden internal leak
  • a list or markdown response that may be safely shaped
  • a shaped response that still contains a forbidden pattern

The negative controls prove that the harness can detect failure. The positive controls prove that it does not reject every useful response.

Step 3 - Follow the Tool Loop

Allow only a bounded number of permitted tool rounds. After each permitted call, continue to the returned text. An unknown tool or a loop that never returns to text remains a failure. Do not silently increase the ceiling to make a test pass.

Step 4 - Score the Final Pipeline

Record metadata rather than private bodies:

scenario_id
tool_rounds
first_completion_kind
final_text_present
initial_verdict
shaping_applied
final_verdict
disposition

The final verdict must be calculated after all shaping. If shaping creates a new sentence boundary or leaves too little useful text, the result should be suppressed.

Step 5 - Compare the Measurement

Run the same scenario set repeatedly and record whether changes affect generation, filtering, shaping, or actual delivery. In the source work, three local runs showed that the raw totals varied, while the more stable result was the reduction in unsafe text reaching the final boundary.

Do not claim a trend from three runs alone. Report the sample size, the stage measured, and the uncertainty.

Security Takeaways

  1. A first completion is an intermediate artifact.
  2. A production-shaped harness must include the context that activates its assertions.
  3. Positive and negative controls are both required.
  4. Shaping is not a bypass; shaped output needs reclassification.
  5. A local owner-controlled canary is safer than testing through an unaware person.

Where This Applies Beyond the Lab

The same pattern applies to email gateways, malware sandboxes, authorization middleware, browser extensions, and data-loss-prevention filters. Whenever a result passes through multiple transformations, test the artifact that crosses the boundary.