π Day 167 β Approval-Gated Automation: Making Privileged Agent Work Accountable
π Topic
Adding an autonomous-supervisor layer that creates a durable task record before sensitive work begins, and gives an operator approval, cancellation, and recovery controls.
π― Goal
Keep low-risk conversational requests simple while treating code changes, files, services, packages, and multi-stage deliveries as supervised work with explicit evidence and operator control.
π What I Did
I implemented a deterministic classifier that leaves only conversational, read-only, non-artifact requests on the direct path. Requests that could modify systems or create deliverables fail closed into supervised handling.
Before the first sensitive action, the supervisor stores the request, goal, success criteria, priority, source channel, and task-scoped permissions. It also records notifications and delivery state so a repeated message or crash does not create duplicate work.
The operator controls are deliberately practical: inspect task status and queue position, approve or reject, cancel safely or immediately, resume, and adjust priority. I also added an emergency CLI and a minimal loopback-only, token-required inbox that read the same durable database rather than depending on the usual chat route.
The evidence stayed scoped to fixtures and a feature flag:
- Phase 7 passed 186 tests, including fresh-process and crash-recovery checks
- Phase 8 expanded the suite to 201 tests and verified unauthenticated inbox access was refused with HTTP 401
- no production database, listener, or LaunchAgent was created during the rollout
- the flag remained unset, so the live communication paths were not silently changed
π Key Cybersecurity Connections
This applies access-control principles to automation. The agent is not trusted because it is helpful; it is constrained because it can affect something valuable.
- least privilege: permissions attach to a task instead of becoming a permanent global ability
- separation of duties: approval and inspection are distinct from execution
- auditability: lifecycle events and notifications have durable records
- safe defaults: uncertain or sensitive requests are supervised, and write actions refuse while the flag is off
- availability: a fallback client can inspect or control work when the primary chat channel is unavailable
π Investigation Questions
- Did the task exist before the first side effect?
- Who approved the action, and what exactly was approved?
- Can a redelivered message create a second task?
- Does the fallback control surface expose more than it needs?
- Can a completed notification be sent twice after a crash?
π¨ Detection Opportunities
- a state-changing task without a durable task record
- an approval command applied to the wrong task or twice
- repeated delivery notifications for the same completed result
- a fallback inbox accepting a request without its required token
- a task that moves to completed without destination confirmation
π§ MITRE ATT&CK Techniques
No direct MITRE ATT&CK mapping claimed. This is a control design exercise for privileged automation and operator oversight.
πΊ Visual Investigation Diagram
Request arrives
β
Read-only conversation or supervised task?
β
Durable task record before action
β
Scoped approval and execution
β
Confirmed result with recoverable audit trail
β Challenges
The classifier has to be conservative without turning every ordinary question into a queue item. The useful boundary was not the wording of a request, but whether it could change files, systems, packages, services, or deliverables.
π What I Learned
The safe version of autonomy is not βthe agent can do anything.β It is βthe agent can make progress inside a well-defined contract, and a person can see, pause, and challenge that progress.β
β‘ Next Steps
- Exercise more edge cases in task classification
- Keep approvals and task permissions easy to inspect
- Test cancellation and resumption under longer-running work
- Preserve the fallback controls as a deliberately narrow emergency path
π§ Reflection
The controls made the automation feel more useful, not less. Knowing that a task has a clear boundary and a recoverable record removes the pressure to trust a black box.
π§© Lessons Learned
What worked
Creating the durable task before the first sensitive side effect.
What broke
The initial temptation was to use the primary chat route as the only control surface.
Why it broke
An operational control should still be reachable when the usual interface is degraded.
Fix / takeaway
Keep a narrow, authenticated fallback that depends on durable task state rather than the main chat channel.
π Skill Progression Context
This is relevant to security engineering and operations because privileged automation needs the same foundations as privileged human access: scope, approvals, logs, recovery, and a way to say no when the evidence is insufficient.
π TL;DR
Sensitive automation now earns a task record, scoped approval, and an audit trail before it earns the right to act.
