🔄 Topic

Hermes needed to write calendar events on my behalf. Instead of pointing it at my real calendar with full write access, I scoped the capability down to a dedicated calendar it owns, banned deletion outright, and used a tombstone pattern for anything that needs to go away.


🎯 Goal

Give an agent real, useful calendar-write capability without letting a bug, a bad model turn, or a misinterpreted instruction destroy data on the calendar I actually rely on.


🛠 What I Did

I scoped the capability to its blast radius before turning it on, and made the scope decision explicit rather than implicit.

Main areas covered:

  • decided Hermes writes only to a dedicated “Hermes” calendar, not to my primary calendar — a new event created by the agent lands somewhere I own and can review, isolated from everything else on my schedule
  • decided the agent can mark-outdated only on events it created itself, never delete — and never touch anything it didn’t create, on any calendar
  • implemented mark-outdated as a tombstone: an event that’s no longer valid gets flagged as outdated rather than removed, so there’s always a record of what the agent decided was stale and why, instead of a silent gap where an event used to be
  • put the enforcement in the executable itself, not in the persona instructions — the constraint holds even if the model’s read of its own instructions drifts, because the code path simply doesn’t expose a delete or cross-calendar-write capability to call
  • documented the two scope decisions and the reasoning explicitly in a handoff, including why enforcement lives in code rather than in the natural-language instruction file — a decision worth being able to point back to later, not just something implied by the diff

🔗 Key Cybersecurity Connections

This is blast-radius containment applied to a personal calendar: instead of asking “can I fully trust this agent with delete access to my real calendar,” the actual question was “what’s the smallest capability that still does the useful thing” — and the answer was write-only, to a dedicated calendar, no deletes, ever. That reframes the trust question from “how much do I trust this system” (which changes over time, in both directions) to “how much could a mistake actually cost” (which is now bounded regardless of how much I trust the system on any given day).

Enforcing the constraint in code rather than in the instruction file is the same lesson as this month’s earlier “narration is not a command” bug: a natural-language instruction is a request the model can misread, drift from, or have overridden by a cleverly-worded input; a capability that’s structurally absent from the executable can’t be talked around. The tombstone pattern is the audit-trail half of the same discipline — even the allowed action (marking outdated) leaves evidence instead of erasing it.


🔍 Investigation Questions

  • Does the agent write to a dedicated, isolated calendar, or directly to the calendar I actually rely on?
  • Is delete capability exposed anywhere in the executable, even if the instructions say not to use it?
  • Does “mark outdated” leave a record, or does it behave like a silent delete under a different name?
  • Is the scope decision documented with its reasoning, or only implicit in what the code happens to do?
  • Would this constraint survive a persona or instruction-file rewrite, since it lives in code rather than in prose?

🚨 Detection Opportunities

Checks for an agent with calendar (or similarly destructive-capable) write access:

  • agent-originated writes landing on a primary or shared calendar instead of an isolated, dedicated one
  • delete capability reachable in code regardless of what the instructions claim
  • a “mark outdated” or similar soft-delete action that doesn’t actually preserve a record
  • a capability scope decision with no documented reasoning to audit later
  • enforcement relying solely on natural-language instructions rather than structural code constraints

Example:

project=hermes-calendar-write
signal=agent_write_capability_reaches_primary_calendar
risk_area=blast_radius_of_agent_originated_data_writes
triage=confirm_writes_isolated_to_dedicated_calendar_delete_unreachable_in_code

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is a data-integrity and least-privilege design pattern for an autonomous agent’s write capability, not an adversary technique.


🗺 Visual Investigation Diagram

Hermes needs calendar write capability
    ↓
Question: full access to primary calendar, or scoped?
    ↓
Decision: dedicated "Hermes" calendar only
    ↓
Decision: mark-outdated, never delete, only on own events
    ↓
Enforcement lives in the executable, not the instruction file
    ↓
Tombstone: outdated events flagged, never erased
    ↓
Blast radius bounded regardless of model behavior on any given day

⚠ Challenges

The tempting shortcut was giving the agent write access to the real calendar directly — it’s less setup, and “it’ll probably behave” felt reasonable given how much other hardening had already gone into the system this month. Choosing the dedicated-calendar-plus-tombstone pattern instead meant more plumbing for a feature that, most days, will never come close to needing the isolation it provides.


📚 What I Learned

I learned that the right question for a new destructive-adjacent capability isn’t “do I trust this system enough,” it’s “what’s the actual cost if I’m wrong about that trust” — and that answering the second question with architecture (isolation, no delete, tombstones) makes the first question matter much less.


➡ Next Steps

  • Periodically review the dedicated Hermes calendar for events that should be reconciled onto the primary one
  • Audit other agent-writable data stores (contacts, reminders) for the same dedicated-store-plus-tombstone pattern
  • Confirm the tombstone flag is actually surfaced somewhere I’d see it, not just stored and forgotten
  • Extend the “enforcement in code, not instructions” principle as a standing checklist item for future write-capable features

🧠 Reflection

This is the kind of security decision that produces literally nothing to show for it on a normal day — no incident, no near-miss, no story. That’s exactly the point: the cost of being wrong about how much to trust the system is now bounded by design, not by how careful the model happens to be that day.


🧩 Lessons Learned

What worked

Scoping writes to a dedicated calendar, banning delete outright, and enforcing both in code rather than in the instruction file.

What broke

Nothing yet — this was proactive containment before the capability shipped, not a response to an incident.

Why it mattered

An agent with real write access to a personal calendar is one misread instruction away from real data loss if the capability isn’t structurally bounded.

Fix / takeaway

Isolate agent-originated writes to a dedicated store, forbid deletion in favor of tombstoning, and enforce the constraint where it can’t be talked around — in code, not in prose.


📈 Skill Progression Context

This supports my cybersecurity progression because blast-radius containment, least-privilege scoping of write capabilities, and structural (rather than policy-only) enforcement are core data-integrity and access-control disciplines for any system with autonomous write access to real user data.


😄 TL;DR

Hermes can write to my calendar now — just not my real one, and never with a delete button, only a tombstone.