Lab Objective

Build a local, no-send confirmation gate that treats speech transcription as untrusted input and allows an agent job to start only after the authorized owner confirms the exact proposal.

Lab Environment

  • Environment: local Python test harness or equivalent small command dispatcher
  • Input types: typed text, voice transcription, and forwarded text
  • Authorized surface: one owner-controlled test channel
  • Safety rule: the harness never starts a real agent job or sends a message

Scenario

A voice note says β€œClaude, check the repository.” Speech recognition may mishear the model name or the command. A forwarded message may look like it came from the owner. The system must offer a proposal and wait for an exact confirmation.

Step 1 - Create the Proposal

Normalize the transcription only enough to display it. Assign a short, expiring confirmation code and store a hash of the proposed action. Do not execute the action at this stage.

Step 2 - Enforce the Caller Boundary

Require the expected platform, owner identifier, private channel, and non-forwarded source. Reject every other fixture before it reaches the dispatcher.

Step 3 - Confirm the Exact Proposal

Accept confirm <code> only from the same authorized principal and only while the code is valid. The confirmation must refer to the stored proposal, not merely to a new piece of text that sounds similar.

Step 4 - Test Failure Fixtures

Test a normal voice proposal, a misheard destructive phrase, a forwarded message, the wrong sender, an expired code, a reused code, and a confirmation for a different proposal. The expected result is offer, reject, reject, reject, reject, reject, and reject respectively.

Security Takeaways

  1. Transcription is data, not authorization.
  2. Confirmation must be bound to the exact proposed action.
  3. The executable gate must enforce identity and channel restrictions.
  4. Tests should prove that destructive-looking or forwarded input cannot start work.
  5. A no-send harness is safer than using a real person as a test endpoint.

Where This Applies Beyond the Lab

The same pattern fits administrative commands, password-reset workflows, deployment approvals, and high-impact actions exposed through chat or voice interfaces.