📅 Day 194 — Stopping the AI Secretary from Writing and Sending Email on Its Own
🔄 Topic
The personal secretary had been asking the Hermes agent to compose email and send it automatically. I stopped both halves of that and replaced them with a three-draft picker over Telegram — real human choice, not a rubber stamp.
🎯 Goal
Make sure no email leaves my accounts without a human actually choosing what gets sent, and give that human something better than one auto-generated draft to approve.
🛠 What I Did
I pulled composition and sending apart, and put a person back in the loop for both.
Main areas covered:
- stopped asking the Hermes agent to write email content directly — composition moved to a controlled path instead of an open-ended generation call
- stopped autosending entirely — no reply leaves an account without an explicit decision
- built a three-draft Telegram picker: instead of one take-it-or-leave-it draft, three real options are generated and a human picks, edits, or rejects
- built the dataset behind the picker so the three drafts are meaningfully different, not three near-duplicates dressed up as choice
- fixed a related bug where outgoing replies did not actually look like email — a formatting gap that would have been a giveaway to anyone receiving one
- logged the model-routing rule for this pipeline as part of the same handoff, so future changes to the underlying model don’t quietly reintroduce autosend-shaped behavior
- guarded against the secretary treating an ordinary sentence as a commitment, while making sure it stopped over-filtering normal English in the process — a false-positive/false-negative pair fixed together
🔗 Key Cybersecurity Connections
Autonomous email composition and sending is one of the highest-consequence capabilities I’ve given any agent — it’s identity, it’s relationships, and it’s often the exact vector used in real business email compromise. Removing autosend is not a downgrade; it’s matching capability to the actual trust I have in the system, which this month’s fail-closed and false-assurance findings have repeatedly shown is lower than a working demo implies.
Giving three genuinely different drafts instead of one is a usability decision with a security payoff: a single draft trains the human to rubber-stamp; a real choice keeps the human actually reading and deciding.
🔍 Investigation Questions
- Can any email leave an account without an explicit human decision?
- Are the three drafts meaningfully different, or cosmetically different?
- Does an outgoing reply look enough like a normal email that a recipient wouldn’t suspect automation?
- Does the commitment-guard now correctly separate real commitments from ordinary conversational English?
- What happens to this pipeline if the underlying model changes?
🚨 Detection Opportunities
Checks for a human-reviewed email pipeline:
- an outgoing message with no corresponding pick-or-approve event
- draft options that are near-identical rather than meaningfully distinct
- outgoing formatting diverging from what a normal client would send
- the commitment guard flagging ordinary sentences or missing real commitments
- a model or routing change landing with no re-verification of the picker’s behavior
Example:
project=hermes-secretary-email
signal=outgoing_message_without_human_decision_event
risk_area=unauthorized_communication
triage=halt_pipeline_confirm_decision_log_matches_every_send
🧭 MITRE ATT&CK Techniques
Possible mapping for the risk being controlled:
- T1534 — Internal Spearphishing (autonomous, unreviewed email generation and sending is the exact shape of this risk turned inward)
🗺 Visual Investigation Diagram
Incoming message needs a reply
↓
Old: agent writes it, agent sends it
↓
New: agent proposes three real drafts
↓
Human picks, edits, or rejects over Telegram
↓
Only a human decision triggers a send
↓
Outgoing reply formatted to look like real email
⚠ Challenges
The commitment-guard fix was fiddlier than expected: too strict and it flags normal conversation as a promise it shouldn’t make; too loose and it misses a real commitment slipping into an autosent draft. Getting both directions right in the same pass took more iteration than the send/autosend fix itself.
📚 What I Learned
I learned that “human in the loop” is only real if the loop offers a real decision. One draft with an approve button is a loop in name only; three meaningfully different drafts is what makes the human’s attention worth anything.
➡ Next Steps
- Watch real usage of the three-draft picker for whether the options stay meaningfully distinct over time
- Keep the model-routing rule under version control alongside the pipeline it protects
- Extend the “looks like real email” formatting check to other outgoing channels
- Revisit the commitment guard’s edge cases as more real traffic passes through it
🧠 Reflection
Removing a capability I had built myself — autosend — felt like a step backward for a moment, until I remembered that the point of this whole month has been matching autonomy to evidence, not to convenience.
🧩 Lessons Learned
What worked
Separating composition from sending, and replacing one auto-approved draft with three real choices.
What broke
Autonomous email writing and sending, and a commitment guard that was miscalibrated in both directions at once.
Why it broke
Convenience had outpaced the actual trust the system had earned.
Fix / takeaway
Give humans a real decision, not a rubber stamp — and match every autonomous capability to evidence of its reliability, not to how impressive the demo looked.
📈 Skill Progression Context
This supports my cybersecurity progression because email-based social engineering and business email compromise are top real-world threats, and building — then correctly constraining — an email-capable agent is direct, hands-on practice in that exact threat model.
😄 TL;DR
The secretary can propose three emails now. It doesn’t get to pick one, and it never gets to hit send.
