📅 Day 165 — Bounded Context, Better Decisions: Testing AI Agent Memory Without Blind Trust
🔄 Topic
Building a bounded, source-attributed working context for a local AI agent, then measuring it without pretending that a benchmark is the same thing as live intelligence.
🎯 Goal
Give an agent the task-relevant facts it needs while preserving provenance, exposing stale or conflicting information, and keeping the existing retrieval path as the default until the evidence supports a change.
🛠 What I Did
I added an opt-in working-wiki layer in front of the existing local retrieval workflow. It compiles a small runtime context from the current goal, constraints, project state, decisions, acceptance gate, relevant guidance, and latest handoff.
The important part was not collecting more text. It was making the selection inspectable. The compiler records source paths, update dates, budgets, and status, and it can flag missing, duplicate, stale, conflicting, and over-budget material.
I then used a deterministic benchmark to compare direct Markdown, the existing retrieval path, and the hybrid approach. The benchmark does not call a model or embeddings, which means its results describe the context-selection process rather than claiming to measure model reasoning.
Evidence from the implementation:
- the new layer is opt-in behind
LOCAL_AGENT_WORKING_WIKI=1 - the existing retrieval path remained the normal default
- the focused working-wiki tests passed 13/13
- the existing retrieval quality check remained above its 0.8 floor at 14/17 (82%)
- a runtime smoke confirmed the generated context and its active acceptance-gate source
🔗 Key Cybersecurity Connections
This is data-integrity work disguised as agent plumbing. An automated system that receives stale instructions, conflicting decisions, or an untraceable summary can make a wrong decision with great confidence.
The defensive controls are familiar:
- provenance: every selected fact has an identifiable source
- integrity checks: conflicts, duplicates, and missing evidence are surfaced instead of silently merged
- least data: the context has a budget, so unrelated history is not automatically exposed to every task
- change control: a new capability is tested behind an explicit switch instead of replacing a working default
🔍 Investigation Questions
- Which source supports this instruction, and how old is it?
- Does a newer handoff contradict a standing rule?
- Is the context missing a required acceptance criterion?
- Did the benchmark measure deterministic selection, or did it quietly make a model-quality claim?
- What happens when the feature flag is not enabled?
🚨 Detection Opportunities
Useful signals for context integrity include:
- a required source is missing from a generated working set
- two active sources disagree on a constraint
- an expired handoff is selected over a newer one
- a context budget is exceeded by low-priority history
- an opt-in component is active when its feature flag is unset
🧭 MITRE ATT&CK Techniques
No direct MITRE ATT&CK mapping claimed. This work is about the integrity and governance of an automated decision-support system.
🗺 Visual Investigation Diagram
Known sources
↓
Select a bounded working set
↓
Preserve provenance and age
↓
Detect stale or conflicting facts
↓
Run the task with explicit context
⚠ Challenges
The tempting shortcut was treating more context as automatically safer. It is not. More unreviewed history can hide the task-critical fact, bring stale instructions back to life, and expose information the task does not need.
📚 What I Learned
More context is not automatically better context. The useful question is whether the agent can show where a fact came from, whether it still applies, and what it displaced. That is how I want to treat any automation that is allowed to influence real work.
➡ Next Steps
- Add more fixture cases for conflicting instructions
- Track which sources are selected most often and why
- Review budget limits against real task types
- Keep the default retrieval path under independent quality checks
🧠 Reflection
This made context feel less like a convenience feature and more like a controlled input channel. The useful question is not whether an agent has a lot to read, but whether its most important facts are current, traceable, and proportionate to the task.
🧩 Lessons Learned
What worked
Keeping the compiler deterministic and source-attributed made the output easy to inspect.
What broke
Early assumptions treated stale material as disposable.
Why it broke
Stale material can still reveal that a newer decision conflicts with an older one.
Fix / takeaway
Keep stale evidence visible to conflict checks, but do not let it silently become an active instruction.
📈 Skill Progression Context
This supports security work because analysts and engineers constantly make decisions from incomplete, aging, and sometimes conflicting evidence. Building systems that label uncertainty and preserve provenance is a practical way to practice that discipline.
😄 TL;DR
Built an opt-in context layer that keeps the useful facts, shows their sources, and calls out the parts that no longer agree.
