🔄 Topic

The personal secretary had been asking the Hermes agent to compose email and send it automatically. I stopped both halves of that and replaced them with a three-draft picker over Telegram — real human choice, not a rubber stamp.


🎯 Goal

Make sure no email leaves my accounts without a human actually choosing what gets sent, and give that human something better than one auto-generated draft to approve.


🛠 What I Did

I pulled composition and sending apart, and put a person back in the loop for both.

Main areas covered:

  • stopped asking the Hermes agent to write email content directly — composition moved to a controlled path instead of an open-ended generation call
  • stopped autosending entirely — no reply leaves an account without an explicit decision
  • built a three-draft Telegram picker: instead of one take-it-or-leave-it draft, three real options are generated and a human picks, edits, or rejects
  • built the dataset behind the picker so the three drafts are meaningfully different, not three near-duplicates dressed up as choice
  • fixed a related bug where outgoing replies did not actually look like email — a formatting gap that would have been a giveaway to anyone receiving one
  • logged the model-routing rule for this pipeline as part of the same handoff, so future changes to the underlying model don’t quietly reintroduce autosend-shaped behavior
  • guarded against the secretary treating an ordinary sentence as a commitment, while making sure it stopped over-filtering normal English in the process — a false-positive/false-negative pair fixed together

🔗 Key Cybersecurity Connections

Autonomous email composition and sending is one of the highest-consequence capabilities I’ve given any agent — it’s identity, it’s relationships, and it’s often the exact vector used in real business email compromise. Removing autosend is not a downgrade; it’s matching capability to the actual trust I have in the system, which this month’s fail-closed and false-assurance findings have repeatedly shown is lower than a working demo implies.

Giving three genuinely different drafts instead of one is a usability decision with a security payoff: a single draft trains the human to rubber-stamp; a real choice keeps the human actually reading and deciding.


🔍 Investigation Questions

  • Can any email leave an account without an explicit human decision?
  • Are the three drafts meaningfully different, or cosmetically different?
  • Does an outgoing reply look enough like a normal email that a recipient wouldn’t suspect automation?
  • Does the commitment-guard now correctly separate real commitments from ordinary conversational English?
  • What happens to this pipeline if the underlying model changes?

🚨 Detection Opportunities

Checks for a human-reviewed email pipeline:

  • an outgoing message with no corresponding pick-or-approve event
  • draft options that are near-identical rather than meaningfully distinct
  • outgoing formatting diverging from what a normal client would send
  • the commitment guard flagging ordinary sentences or missing real commitments
  • a model or routing change landing with no re-verification of the picker’s behavior

Example:

project=hermes-secretary-email
signal=outgoing_message_without_human_decision_event
risk_area=unauthorized_communication
triage=halt_pipeline_confirm_decision_log_matches_every_send

🧭 MITRE ATT&CK Techniques

Possible mapping for the risk being controlled:

  • T1534 — Internal Spearphishing (autonomous, unreviewed email generation and sending is the exact shape of this risk turned inward)

🗺 Visual Investigation Diagram

Incoming message needs a reply
    ↓
Old: agent writes it, agent sends it
    ↓
New: agent proposes three real drafts
    ↓
Human picks, edits, or rejects over Telegram
    ↓
Only a human decision triggers a send
    ↓
Outgoing reply formatted to look like real email

⚠ Challenges

The commitment-guard fix was fiddlier than expected: too strict and it flags normal conversation as a promise it shouldn’t make; too loose and it misses a real commitment slipping into an autosent draft. Getting both directions right in the same pass took more iteration than the send/autosend fix itself.


📚 What I Learned

I learned that “human in the loop” is only real if the loop offers a real decision. One draft with an approve button is a loop in name only; three meaningfully different drafts is what makes the human’s attention worth anything.


➡ Next Steps

  • Watch real usage of the three-draft picker for whether the options stay meaningfully distinct over time
  • Keep the model-routing rule under version control alongside the pipeline it protects
  • Extend the “looks like real email” formatting check to other outgoing channels
  • Revisit the commitment guard’s edge cases as more real traffic passes through it

🧠 Reflection

Removing a capability I had built myself — autosend — felt like a step backward for a moment, until I remembered that the point of this whole month has been matching autonomy to evidence, not to convenience.


🧩 Lessons Learned

What worked

Separating composition from sending, and replacing one auto-approved draft with three real choices.

What broke

Autonomous email writing and sending, and a commitment guard that was miscalibrated in both directions at once.

Why it broke

Convenience had outpaced the actual trust the system had earned.

Fix / takeaway

Give humans a real decision, not a rubber stamp — and match every autonomous capability to evidence of its reliability, not to how impressive the demo looked.


📈 Skill Progression Context

This supports my cybersecurity progression because email-based social engineering and business email compromise are top real-world threats, and building — then correctly constraining — an email-capable agent is direct, hands-on practice in that exact threat model.


😄 TL;DR

The secretary can propose three emails now. It doesn’t get to pick one, and it never gets to hit send.