🧪 Lab 08 – Auditing What an AI Agent Is Given to Read Before It Acts
Lab Objective
Practice treating an AI agent’s task context as security-relevant input rather than an invisible prompt. The objectives were to:
- compile a small task-specific working set from known sources
- preserve the origin and age of each selected fact
- surface missing, stale, duplicate, conflicting, and over-budget information
- compare selection strategies deterministically
- keep the feature opt-in while the existing path remains the default
Lab Environment
- System: local macOS workstation
- Component: working-wiki compiler for a local agent
- Inputs: goal, constraints, project state, decisions, acceptance gate, guidance, and latest handoff
- Tools: Python, Markdown fixtures,
rg, SHA-256 checks, focused tests, deterministic benchmark - Safety boundary: the context layer was enabled only with
LOCAL_AGENT_WORKING_WIKI=1
Scenario
An agent needs enough context to complete a task, but the repository contains old handoffs, duplicate decisions, conflicting notes, and unrelated history. Feeding every file to the model is both inefficient and risky. A compact context must be useful, traceable, and honest about uncertainty.
Step 1 - Inventory the Candidate Sources
Start by identifying the documents that could influence a task.
rg --files docs handoff | sort
rg -n -i "acceptance|constraint|decision|rollback" docs handoff
The point is not to select every hit. It is to identify the sources that govern the current task.
Step 2 - Record Provenance Before Summarizing
For each selected source, record its path, modification date, and digest.
stat -f "%Sm %N" -t "%Y-%m-%d %H:%M" handoff/latest-task.md
shasum -a 256 handoff/latest-task.md
An agent summary without a source is difficult to challenge later. Provenance makes a fact inspectable.
Step 3 - Make Contradictions Visible
Search for competing instructions before the context is assembled.
rg -n -i "default.*enabled|default.*disabled|current gate" docs handoff
In the real implementation, stale sources were not simply hidden. They remained visible to the conflict detector, because an old instruction can be exactly what explains a contradiction.
Step 4 - Run the Deterministic Checks
The focused test and benchmark exercised source selection without asking a model to judge itself.
python3 AI-OS/tests/test_working_wiki.py
python3 scripts/working-wiki-benchmark.py
python3 scripts/rag-eval.py --threshold 0.8
The completed run passed 13 focused tests. The existing retrieval quality check remained above its configured threshold at 14/17 (82%).
Step 5 - Keep the New Path Opt-In
Run the feature only when explicitly requested.
LOCAL_AGENT_WORKING_WIKI=1 scripts/local-agent --help
Leaving the flag unset preserves the existing behavior. This is a deployment control, not merely a convenience switch.
Security Takeaways
- Context is input. Old or conflicting task data can be as dangerous as an unvalidated configuration file.
- Provenance makes review possible. A source path and digest let a reviewer inspect the claim instead of trusting a summary.
- Stale data is evidence too. It may be wrong for action, but still useful for detecting conflict.
- Budgets reduce exposure. The task does not need every historical note.
- Benchmarks need honest labels. A deterministic selection benchmark measures selection, not model intelligence.
Where This Applies Beyond AI
The same workflow applies to incident timelines, threat-intelligence enrichment, compliance evidence, and configuration baselines. Before acting, identify the source, verify that it still applies, find contradictions, and keep the evidence small enough to review.
