๐ Day 188 โ Giving Agent Memory Real Sources, and Routing Trivial Turns to a Small Model
๐ Topic
Two upgrades to how Hermes thinks day to day: a memory index that tracks real sources and supersession instead of an undifferentiated pile of facts, and a tiering proxy that stops routing trivial turns to an expensive model.
๐ฏ Goal
Make the assistantโs memory trustworthy enough to cite, and make its resource use match the actual difficulty of what it is being asked.
๐ What I Did
I upgraded memory quality and cost discipline together.
Main areas covered:
- gave the Hermes memory index real sources: every remembered fact now carries where it came from and can be retrieved with that provenance intact, not just as bare text
- added supersession ranking, so when a newer fact contradicts or updates an older one, retrieval knows which one is current instead of surfacing both with equal weight
- added write-back so memory updates persist properly instead of living only in a sessionโs working state
- built a dedicated recall skill so retrieving from memory is a first-class, auditable action rather than an implicit side effect of a prompt
- added a model-tiering proxy that routes trivial turns โ acknowledgments, short confirmations, simple lookups โ to a small model, reserving the larger model for turns that actually need it
- documented both changes as a paired handoff, since better memory and cheaper routing reinforce each other: a system that remembers correctly needs fewer expensive turns to re-derive context it should already have
๐ Key Cybersecurity Connections
Provenance in a memory system is the same principle as provenance in a change-control system: a fact worth acting on should be traceable to where it came from, and a superseded fact should not silently outrank the correction that replaced it. An assistant that cannot distinguish current from stale information is making decisions on data it should have already discarded.
The tiering proxy is a cost and exposure decision as much as a performance one: fewer turns sent to the larger model means less context handed to it by default, which matters for the same reasons token budgeting mattered a few weeks ago.
๐ Investigation Questions
- Can every remembered fact be traced to its source?
- When two facts conflict, does retrieval know which one is current?
- Does memory persist correctly, or only within one session?
- Is recall an auditable action, or an invisible side effect?
- Does the tiering proxy ever misroute a turn that actually needed the larger model?
๐จ Detection Opportunities
Checks for a memory and routing upgrade:
- retrieved fact with no source attribution
- superseded fact still surfacing above its replacement
- memory update that does not survive past the current session
- a turn misrouted to the small model despite needing real reasoning
- recall skill invoked without a corresponding log entry
Example:
project=hermes-memory-and-tiering
signal=superseded_fact_outranking_current_one
risk_area=stale_data_driving_decisions
triage=inspect_supersession_ranking_and_source_provenance
๐งญ MITRE ATT&CK Techniques
No direct mapping claimed. This is data-quality and resource-governance work for a personal assistantโs memory layer.
๐บ Visual Investigation Diagram
Fact learned
โ
Source + provenance recorded
โ
Supersession ranking on conflict
โ
Write-back persists it for real
โ
Recall skill retrieves it, auditable
โ
Trivial turns โ small model; real turns โ full model
โ Challenges
The subtle risk in supersession ranking is getting it backwards โ trusting recency when the older fact was actually the correct one and the newer one was a mistaken correction. Ranking by time alone is a heuristic, not a guarantee, and it is worth remembering which one it is.
๐ What I Learned
I learned that memory quality and cost efficiency are not separate concerns here โ an assistant with accurate, well-sourced memory needs less expensive re-reasoning to recover context it should already hold, which makes the tiering savings real rather than just cheaper wrong answers.
โก Next Steps
- Watch for cases where recency ranking overrides a still-correct older fact
- Audit the tiering proxyโs misroute rate on real traffic
- Keep provenance mandatory for anything written back to memory
- Extend the recall skillโs audit trail to cover partial or low-confidence retrievals
๐ง Reflection
Building memory as something with sources and ranking, rather than a bag of facts, changed how much I trust what the assistant tells me it โremembersโ โ which was the actual point.
๐งฉ Lessons Learned
What worked
Pairing provenance-backed memory with cost-aware routing in the same session.
What broke
Nothing catastrophic, but the recency-as-truth heuristic in supersession ranking needs watching.
Why it matters
An assistantโs memory should be trustworthy enough to cite, and its resource use should match the actual difficulty of the task.
Fix / takeaway
Provenance and supersession first, cheaper routing second โ cost savings built on unreliable memory are savings on the wrong thing.
๐ Skill Progression Context
This supports my cybersecurity progression because provenance, data currency, and resource-appropriate access are recurring themes across identity systems, logging, and cost-governed infrastructure.
๐ TL;DR
Gave the assistantโs memory real sources and a sense of which fact is current โ and stopped burning a big model on โok.โ
