πŸ”„ Topic

Turning a pile of local AI tools into an β€œoperating system”: routing, governance rules, task contracts, and verified execution.


🎯 Goal

Make agent autonomy safe and repeatable instead of ad hoc: every task routed, every capability governed, every result verified.


πŸ›  What I Did

I built up the AI-OS layer of my local agent workstation. Durable context lives in structured files β€” docs, handoffs, memory, playbooks, skills β€” so no single model has to remember everything and any tool can pick up the work. I added a routing orchestrator that classifies tasks and escalates hard ones to a stronger model, controlled execution hooks, retry loops with temperature and prompt variation, a self-review pipeline stage, and a Task Contract + Verification Gateway so autonomous work has explicit scope and checked results. The system also routes agent outputs into my Obsidian vault: smoke tests confirmed audit notes and drafts land in the right inbox folders automatically.

Main areas covered:

  • structured file-based memory and handoffs
  • routing orchestration with escalation to stronger models
  • controlled execution hooks
  • task contracts and a verification gateway
  • self-review as a pipeline stage
  • routed capture into the Obsidian vault
  • written governance rules for agents in the repo

πŸ”— Key Cybersecurity Connections

This is security architecture applied to AI: least privilege via governance rules, separation of duties via the verification gateway, and audit trails via handoffs and routed notes. An autonomous agent without contracts and verification is an unaccountable privileged process.


πŸ” Investigation Questions

  • What is each agent allowed to do, and where is that written?
  • Does every autonomous task have a defined scope and success check?
  • When does the router escalate instead of letting a weak model guess?
  • Are agent outputs captured somewhere reviewable?
  • What prevents the agent from storing secrets in its memory files?

🚨 Detection Opportunities

Potential monitoring ideas:

  • tasks executed without a matching contract
  • verification gateway failures
  • escalations spiking for a task category
  • agent writes outside approved capture paths
  • governance file modifications

Example:

project=ai-os-workstation
change_type=autonomous_task_execution
risk_area=unverified_agent_output
triage=check_contract_scope_and_gateway_result

🧭 MITRE ATT&CK Techniques

Possible mappings depending on confirmed behavior:

  • T1059 β€” Command and Scripting Interpreter
  • T1078 β€” Valid Accounts
  • T1565 β€” Data Manipulation

πŸ—Ί Visual Investigation Diagram

Incoming task
    ↓ Router / classifier
    ↓ Task contract
    ↓ Local model or escalation
    ↓ Verification gateway
    ↓ Routed, reviewable output

⚠ Challenges

The challenge was resisting elaborate systems before there is a repeated need. Governance that nobody follows is worse than none, so the rules had to stay small, written, and enforced by the pipeline itself.


πŸ“š What I Learned

I learned that trust in automation is built from boring parts: written rules, explicit contracts, verification steps, and logs. The same ingredients that make a SOC trustworthy make an agent trustworthy.


➑ Next Steps

  • Benchmark the routing decisions against real tasks
  • Keep secrets categorically out of memory files
  • Expand playbooks only when workflows actually repeat
  • Review escalation logs weekly

🧠 Reflection

This was useful because designing controls for my own agents made abstract security principles concrete: I am the admin, the auditor, and the user of this tiny privileged system.


🧩 Lessons Learned

What worked

Task contracts plus a verification gateway before more autonomy.

What broke

Early sessions where a weak local model confidently produced wrong work.

Why it broke

No routing or verification meant nothing caught the failure.

Fix / takeaway

Route by difficulty, escalate when needed, and never skip verification.


πŸ“ˆ Skill Progression Context

This supports my cybersecurity progression because governing autonomous systems β€” scoping, verifying, logging β€” is the same discipline as securing any privileged automation, and it is becoming a real job skill.


πŸ˜„ TL;DR

Gave my agents a constitution and a border control.