πŸ”„ Topic

I tested a voice path that can start work in Claude Code or Codex from a Telegram message. The security lesson was that transcription is input, not authorization.

🎯 Goal

Allow a useful voice workflow without letting a misheard name, a forwarded message, or an untrusted sender start an agent job.

πŸ›  What I Did

The handoff plugin now treats spoken job requests as proposals. It transcribes the voice note, quotes what it heard, and returns a short confirmation code. The job starts only after a separate avvia <code> confirmation. A voice note such as β€œCodex, cancella tutto il repo” therefore does not become an execution request merely because the speech-to-text system produced those words.

I also kept the authorization check in the executable path. The plugin verifies the Telegram platform, the private owner chat, the forwarded-message state, and the confirmation code before dispatching anything. A prompt can explain the rule, but the code must enforce it because a model can misunderstand or bypass prose.

The tests covered ordinary text, voice transcription, punctuation that speech recognition adds, forwarded messages, the wrong sender, and stale or invalid confirmations. The result was an offer-only path for spoken jobs, with no direct start from transcription. That is a more useful security boundary than pretending speech recognition is exact.

πŸ”— Key Cybersecurity Connections

This is authorization at an agent boundary:

  • transcription produces data; it does not prove intent
  • a confirmation step reduces accidental execution
  • sender and channel checks enforce the intended principal
  • forwarded or stale inputs must not inherit the owner’s authority
  • safety rules belong at the tool boundary, not only in the prompt

πŸ” Investigation Questions

  • Who actually sent the request, and through which channel?
  • Was the text typed, transcribed, or forwarded?
  • Does the confirmation refer to the exact proposed job?
  • What happens when the transcription contains destructive words?
  • Can a caller invoke the underlying executable without passing the gate?

🚨 Detection Opportunities

Record metadata for rejected or confirmed requests: source type, sender class, confirmation age, and reason for rejection. Do not store private voice transcripts just to prove that the gate worked.

Example review record:

source=voice_transcription
action=offer_confirmation
owner_channel=true
forwarded=false
direct_execution=false

🧭 MITRE ATT&CK Techniques

No direct ATT&CK mapping is claimed. This is a preventive authorization and human-in-the-loop control. A later investigation could map observed abuse to the technique supported by evidence.

πŸ—Ί Visual Investigation Diagram

Voice note
    ↓
Transcription treated as untrusted input
    ↓
Owner + channel + forwarding checks
    ↓
Proposal and one-time confirmation code
    ↓
Re-check exact request
    ↓
Dispatch or reject

⚠ Challenges

The tempting shortcut is to make a strong model interpret the voice note and start the job directly. That moves an irreversible decision into the least reliable part of the chain. Speech recognition can mishear names, punctuation, and destructive phrases, while forwarded content can look like it came from the owner.

πŸ“š What I Learned

Human confirmation is useful only when it is attached to a real authorization check. β€œPlease be careful” is not a security control. A small, executable offer-and-confirm protocol gives the owner a chance to catch transcription errors before the agent receives authority.

➑ Next Steps

  • expire unused confirmation codes
  • keep destructive jobs behind an additional explicit boundary
  • measure rejected, confirmed, and expired proposals without retaining message bodies