๐Ÿ”„ Topic

The memory-fabricโ€™s citation verifier had a naming problem masking a real one: content that hash-verified correctly could still be wrongly treated as trustworthy, because hash integrity and promotion eligibility had been conflated into one flag.


๐ŸŽฏ Goal

Make โ€œthis citationโ€™s bytes match its hashโ€ and โ€œthis citation is safe to promoteโ€ two separate, honestly-named questions โ€” and close the path-handling gaps hiding between them.


๐Ÿ›  What I Did

I tightened the citation pipeline at both the naming and the logic level.

Main areas covered:

  • rejected absolute citation locations explicitly, before any path join โ€” previously a location that happened to resolve beneath root by coincidence could slip through
  • confirmed traversal and symlink escapes were already caught by resolve-then-relative-to logic, and added explicit tests proving it rather than assuming it
  • fixed duplicate handling: a location repeated within one manifest now resolves deterministically, first occurrence wins, every later duplicate reported as failed โ€” instead of silently re-verifying or letting a later entry ride on an earlier oneโ€™s verified status
  • found the promotion loop had the identical latent gap โ€” set membership does not distinguish a first occurrence from a second one under a different key โ€” and mirrored the same first-occurrence tracking there
  • renamed IndexRebuildResult.verified to hash_verified (keeping a compatibility alias) to make explicit that it only ever meant hash-integrity success, not full promotion eligibility
  • confirmed evaluate_promotion remains the sole authoritative gate โ€” a location can hash-verify and still be refused for being a sensitive location, matching a secret pattern, or being a duplicate

๐Ÿ”— Key Cybersecurity Connections

A hash proves bytes have not changed since it was computed. It says nothing about whether those bytes should be trusted, whether the path pointed somewhere it should not, or whether this is the second copy of something that already appeared once. Naming a field verified when it only means โ€œhashes matchedโ€ invites every future caller to over-trust it.

The duplicate-handling fix is a data-integrity classic: without deterministic first-occurrence tracking, the same manifest processed twice โ€” or a manifest with the same location listed under two keys โ€” could non-deterministically produce different trust outcomes for identical input.


๐Ÿ” Investigation Questions

  • Does hash verification get treated as sufficient for trust anywhere it should not be?
  • Are absolute paths rejected before any path-joining logic runs, or after?
  • What happens when one location appears twice in a manifest, under the same or different keys?
  • Is there exactly one authoritative gate for promotion, or could logic drift between two?
  • Do field names describe what they actually guarantee?

๐Ÿšจ Detection Opportunities

Checks for a citation or manifest integrity pipeline:

  • a verified or similarly named field consumed as a full trust signal downstream
  • absolute path accepted into join logic before an explicit rejection check
  • duplicate manifest entries producing non-deterministic verification outcomes
  • promotion decision made anywhere other than the single authoritative gate
  • sensitive-location or secret-pattern content passing because its hash matched

Example:

project=memory-fabric
signal=hash_verified_field_treated_as_full_trust
risk_area=integrity_vs_authorization_conflation
triage=trace_field_name_to_every_consumer_and_check_promotion_gate

๐Ÿงญ MITRE ATT&CK Techniques

Possible mappings for the risk class:

  • T1565 โ€” Data Manipulation
  • T1006 โ€” Direct Volume Access (path traversal is a cousin risk in the same family)

๐Ÿ—บ Visual Investigation Diagram

Citation location + content
    โ†“
Reject absolute paths explicitly
    โ†“
First occurrence wins; duplicates fail
    โ†“
Hash verified? (integrity only)
    โ†“
evaluate_promotion: sensitive? secret pattern? duplicate?
    โ†“
Promotion โ€” the only real trust decision

โš  Challenges

The renaming felt like the smallest change in the whole fix and was arguably the most important one. verified sounds finished; hash_verified sounds like exactly what it is โ€” one check among several, not the conclusion.


๐Ÿ“š What I Learned

I learned that field names are part of the security surface. A well-implemented check with a misleadingly complete-sounding name will eventually get trusted for more than it proves, by someone who never read the implementation.


โžก Next Steps

  • Audit other boolean fields in the stack for similar naming-versus-guarantee gaps
  • Keep evaluate_promotion as the only place promotion decisions are made
  • Add manifest fuzzing with deliberately duplicated and absolute-path entries
  • Revisit this pattern whenever a new content-ingest pipeline is added

๐Ÿง  Reflection

This was a reminder that integrity and authorization are different questions, and code that answers one should never be named as if it answered the other.


๐Ÿงฉ Lessons Learned

What worked

Separating hash-integrity from promotion eligibility, both in logic and in the fieldโ€™s name.

What broke

Absolute-path handling, duplicate-location handling, and a field name that overpromised.

Why it broke

A single flag was asked to mean two different things, and path handling had gaps a coincidental โ€œunder rootโ€ resolution could hide.

Fix / takeaway

Name fields for exactly what they guarantee, keep one authoritative decision point, and handle duplicates deterministically.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because conflating integrity with authorization is a recurring real-world flaw, and path-traversal and duplicate-handling discipline are core data-integrity skills.


๐Ÿ˜„ TL;DR

A hash matching is not the same as a citation being trustworthy โ€” now the code says so.