📅 Day 158 — Gated Autonomy: A Phone Approval Loop Before Any Agent Sends Anything
🔄 Topic
I built an approval pipeline so autonomous outreach can research and draft on its own, but nothing gets sent until I approve it from my phone.
🎯 Goal
Let the agent do the work without letting it act externally: every send gated behind an explicit human approval, over channels that fail safe.
🛠 What I Did
I wired a human-in-the-loop layer between the agent and the outside world.
Main areas covered:
- built a gated outreach run: research and draft happen autonomously, then execution stops at a phone gate
- set up a working Telegram approval loop, built, activated, and live-tested end to end
- added an n8n approval layer, found its IMAP trigger flaky, and replaced it with a host-side email-approval poller plus a watchdog
- hardened the email-reply approval channel to accept only [id:]-tagged replies, so a busy inbox cannot accidentally approve anything
- emitted drafts as JSON through a bridge so the approval hook sees exactly what would be sent
- added an outreach exclusion guard so certain recipients can never be contacted automatically
- verified the first real send end to end through the host approval hook
🔗 Key Cybersecurity Connections
An agent that can email strangers on my behalf is an impersonation engine with my name on it. The approval gate is the control that keeps it a tool. The design choices are classic security engineering: explicit allow (the [id:]-only rule), fail-safe defaults (flaky trigger replaced, watchdog added), deny lists (exclusion guard), and non-repudiation (I approved exactly the JSON draft that got sent).
🔍 Investigation Questions
- Can anything reach an external recipient without passing the gate?
- What happens when the approval channel is down — queue, or fail open?
- Could a random email in my inbox be misread as an approval?
- Is what I approve byte-identical to what gets sent?
- Who can never be contacted, and is that list enforced in code?
🚨 Detection Opportunities
Checks for a gated-autonomy pipeline:
- send executed with no matching approval record
- approval accepted from a reply without a valid [id:] tag
- approval poller silent longer than its watchdog window
- draft modified between approval and send
- excluded recipient appearing in any draft queue
Example:
project=outreach-approval-pipeline
signal=send_without_matching_approval_id
risk_area=unauthorized_external_action
triage=halt_pipeline_compare_sent_payload_to_approved_json
🧭 MITRE ATT&CK Techniques
Possible mappings if this pipeline were abused:
- T1534 — Internal Spearphishing
- T1585 — Establish Accounts
The controls exist precisely so the pipeline cannot become these.
🗺 Visual Investigation Diagram
Agent researches + drafts
↓
JSON draft to approval hook
↓
Phone gate (Telegram / [id:] email reply)
↓
Approved? → send exactly that payload
↓
Not approved / channel down → nothing happens
⚠ Challenges
The flaky n8n IMAP trigger was the lesson: an unreliable approval channel is worse than none, because it trains you to bypass it. Replacing it with a boring host poller plus a watchdog made the gate something I can actually trust.
📚 What I Learned
I learned that the approval message format is a security boundary. Before the [id:]-only rule, any casual reply in my inbox was one parsing bug away from being an approval.
➡ Next Steps
- Keep the exclusion guard list reviewed and versioned
- Log every approval decision with the draft hash
- Extend the same gate pattern to other external actions
- Periodically test the fail-closed behavior on purpose
🧠 Reflection
This was the most security-shaped building I have done: the feature is a restriction. The pipeline is valuable because of what it refuses to do without me.
🧩 Lessons Learned
What worked
Fail-closed design: no approval, no send, no exceptions.
What broke
The n8n IMAP trigger silently missing approval emails.
Why it broke
A third-party trigger was a single, unmonitored point of failure on a security control.
Fix / takeaway
Controls need watchdogs. A gate nobody monitors quietly becomes a decoration.
📈 Skill Progression Context
This supports my cybersecurity progression because designing authorization boundaries, fail-safe channels, and audit trails around automation is the same discipline as building any privileged workflow control.
😄 TL;DR
The agent drafts, my phone decides, and nothing sends itself.
