📅 Day 199 — Investigating an Unexplained Mac Shutdown and Restricting Who Can Power It Off
🔄 Topic
A WhatsApp message powered my Mac off. Nothing in the agent’s own turn could have caused it — no model response had even landed yet, no tool had executed — which meant the real caller was unidentified and unaccountable. Fixing that took a default-deny source gate on the shutdown script, a real fix to the tier-proxy path that let it happen, a repaired halt sequence, and — a day later — nearly losing both fixes in a 70-commit branch merge.
🎯 Goal
Make sure an emergency shutdown can only ever be triggered by an attributable, allow-listed source, and confirm that fix actually survives contact with the rest of the codebase.
🛠 What I Did
I traced an unattributable shutdown back to its enabling condition, closed it structurally, then had to defend the fix during a large merge.
Main areas covered:
- investigated a Mac power-off triggered by a WhatsApp message: the agent turn provably could not have caused it — no model response yet, no tool execution, a single tool schema in play — so the actual caller was never identified
- fixed the tier-proxy behavior that made an emergency shutdown reachable from an impersonation surface in the first place, rather than only patching the symptom downstream
- added a default-deny source gate to
power-off.sh: it now requires--source=<name>from an explicit allowlist, and records full caller process ancestry on every attempt, so an unattributed invocation becomes attributable the next time instead of vanishing - scoped
HERMES_CONTEXT.mdto Telegram/CLI only, narrowing where a shutdown-capable context can legitimately originate - explicitly documented what the fix does not close: the sudoers path still lets anything call
/sbin/shutdowndirectly — narrowing that needs a root-owned wrapper and my own password, which is future work, not a false sense of done - separately repaired the halt ladder itself: a dead AppleScript rung had failed 3/3 recorded attempts (Apple Events not permitted) while burning 25 seconds, and a fixed 20-second wait was too short, causing the script to escalate and kill apps needlessly — replaced with polled 45s/45s/60s waits, sudoers scoped to exactly
/sbin/shutdown -h, and the script versioned for the first time so a missing checkout can never quietly mean “laptop never shuts down” - a day later, merging a long-diverged Hermes/WhatsApp branch back into main, discovered both files — the power-off source gate and the tier-proxy impersonation fix — existed only on the branch; main’s copy of the tier proxy still had the pre-fix behavior, so redeploying from main would have silently re-armed shutdown-by-WhatsApp-message
- resolved twelve merge conflicts file-by-file based on which side had actually kept developing, rather than a blanket preference for one branch, and verified 788 tests passing with no conflict markers left in the tree before trusting the merge
🔗 Key Cybersecurity Connections
An emergency shutdown is a destructive, high-consequence action, and “the agent’s own turn couldn’t have done it” is exactly the kind of finding that should escalate an investigation, not close it — an unattributed trigger with no identified caller is worse than a known bug, because it means the authorization boundary itself has a gap. Default-deny on a destructive action’s source is the correct shape of fix: allow-list who may call it, log ancestry on every attempt, and say plainly what’s still open rather than declaring victory.
The merge near-miss is its own lesson: a security fix that exists only on an unmerged branch is not actually deployed. “I fixed it” and “the fix is reachable from what actually ships” are two different claims, and this month’s recurring false-assurance pattern would have caught this one too if the merge had gone through as a blanket main-wins resolution.
🔍 Investigation Questions
- Is every destructive or irreversible action gated to an explicit, allow-listed set of callers?
- Does a security fix live only on a branch, or is it actually reachable from what’s deployed?
- When an incident’s proximate cause can’t be identified, does the response escalate scope or just patch the visible symptom?
- Does a safety mechanism’s own escalation ladder (retries, waits, fallbacks) actually work, or does it just look complete on paper?
- What does the fix explicitly not cover, and is that gap documented rather than implied to be closed?
🚨 Detection Opportunities
Checks for destructive-action authorization:
- a destructive or emergency action reachable without an explicit, allow-listed source check
- a security fix present on a feature branch but absent from what’s actually deployed from main
- an unattributed trigger for a high-consequence action, closed without identifying or logging the real caller
- a safety ladder’s fallback step silently failing (e.g., 3/3 recorded failures) without anyone noticing until it’s needed
- a fix described as “done” without an explicit list of what it does not yet close
Example:
project=hermes-power-off-gate
signal=destructive_action_reachable_without_source_allowlist
risk_area=unattributed_high_consequence_trigger
triage=require_explicit_source_arg_log_caller_ancestry_before_execution
🧭 MITRE ATT&CK Techniques
Possible mapping for the risk being controlled:
- T1531 — Account Access Removal / service disruption via an unattributed, unauthorized destructive command (closest adjacent framing; this is an access-control gap on a locally-triggered destructive action rather than a named intrusion technique)
🗺 Visual Investigation Diagram
WhatsApp message triggers a real shutdown
↓
Agent turn provably innocent — no response, no tool exec
↓
Real caller unidentified
↓
Fix tier-proxy: impersonation surface can't reach shutdown
↓
Default-deny source gate + caller ancestry logging on power-off.sh
↓
Repair the dead halt-ladder rung, poll waits instead of fixed timeout
↓
One day later: branch merge nearly drops both fixes from main
↓
Resolve per-file, verify 788 tests, confirm veto intact in the merged tree
⚠ Challenges
The scariest part wasn’t the bug — it was the honest admission that the real caller was never identified, only that the agent turn couldn’t have been it. Shipping a fix without full attribution meant explicitly documenting the remaining gap (the sudoers path) instead of letting the allowlist fix imply more coverage than it actually has. The merge conflict a day later was its own scare: losing the fix silently, by picking the wrong side of a conflict, would have been worse than never finding the bug at all — it would have looked fixed.
📚 What I Learned
I learned that default-deny belongs on any action with real-world, irreversible consequences — not as a philosophy, but as a specific requirement: an explicit allowlist, logged attribution, and a documented boundary of what’s still open. I also learned that a fix isn’t real until it’s confirmed reachable from what’s actually deployed, which is a distinct check from “the diff looks correct.”
➡ Next Steps
- Close the remaining sudoers gap with a root-owned wrapper, once I’m ready to spend the setup time
- Add a regression test asserting the tier-proxy impersonation veto specifically, so it can’t silently regress in a future merge
- Extend caller-ancestry logging to other destructive scripts beyond power-off
- Schedule a periodic check that branch-only security fixes get merged promptly, not left drifting
🧠 Reflection
Finding a bug you can’t fully explain is uncomfortable, but shipping a fix without pretending you’ve fully explained it turned out to be the right call — the allowlist and logging close the practical gap even without a named culprit, and the documented “does NOT close” line kept me honest about what remained.
🧩 Lessons Learned
What worked
Default-deny source gating plus caller-ancestry logging, and reviewing the merge conflict file-by-file instead of trusting a blanket resolution strategy.
What broke
An emergency shutdown was reachable from an impersonation surface with no attributable caller, the halt ladder’s GUI fallback silently failed every time it was tried, and both fixes nearly vanished from main during a large branch merge.
Why it mattered
Destructive, irreversible actions need the highest bar for authorization, and a fix that only exists on an unmerged branch provides zero real protection.
Fix / takeaway
Gate destructive actions to an explicit allowlist with logged attribution, document what’s still open rather than implying full closure, and verify a security fix survives merges into what’s actually deployed.
📈 Skill Progression Context
This supports my cybersecurity progression because authorization design for destructive actions, default-deny gating, and confirming a fix is actually deployed rather than merely committed are core access-control and change-management disciplines in real security engineering.
😄 TL;DR
Something powered my Mac off over WhatsApp with no identifiable agent action behind it — so I default-denied who’s allowed to pull the plug, fixed the broken halt ladder, and then nearly lost both fixes in a merge the very next day.
