📅 Day 144 — A 348-Term Glossary and the False Positives That Came With It
🔄 Topic
Building a broad cybersecurity, networking, Linux, and developer-tooling glossary across my study vault, then linking every matching mention — and cleaning up the collateral damage that caused.
🎯 Goal
Turn scattered vocabulary across months of notes into a single, sourced reference layer, without corrupting the notes it touches.
🛠 What I Did
I scanned the vault for cybersecurity, Linux, shell, networking, web security, SOC, and adversary-behavior vocabulary, expanded a seed list into 348 terms, and generated one sourced note per term with a plain explanation, a why-it-matters line, and backlinks to where it appeared. Then I linked every matching mention across existing notes back into the glossary.
Main areas covered:
- generating 349 files: 348 term notes plus one index
- scanning 485 source files and updating 410 of them with new links
- inserting 3,820 links across the vault
- sourcing terms from NIST CSRC, MDN, OWASP, MITRE ATT&CK, RFCs, and official Git/GitHub Pages docs
- validating every generated and modified file for broken or unbalanced links
- grouping the 348 terms into five topic maps: Networking, Linux Privilege Escalation, Web Security, SOC Operations, and Developer Tooling
🔗 Key Cybersecurity Connections
Mass find-and-link automation across hundreds of files is exactly the kind of bulk change that needs a backup and a validation pass before you trust it — the same reasoning that applies to any bulk remediation script run against production data. I took a pre-link backup archive before running anything, which mattered.
🔍 Investigation Questions
- Did the linker touch anything it shouldn’t have — frontmatter, code blocks, URLs, existing links?
- What do the false positives look like, and why did they happen?
- Is every generated link’s target actually present, or does it point at nothing?
- Can the change be fully reverted from the backup if needed?
🚨 Detection Opportunities
Potential monitoring ideas:
- bulk-edit jobs that modify hundreds of files without a prior backup
- generated links whose target file doesn’t exist (broken reference)
- automated text-matching that fires inside code blocks or raw syntax it shouldn’t touch
- content restored from backup after an automated pass — a signal the pass needs tighter rules
Example:
project=cybersecurity-glossary
change_type=bulk_autolink_pass
risk_area=unintended_content_modification
triage=validate_then_diff_against_backup
🧭 MITRE ATT&CK Techniques
Not directly applicable to this task — this was a knowledge-base build, not adversary behavior. No mappings claimed here.
🗺 Visual Investigation Diagram
Scan vault for vocabulary
↓
Generate sourced term notes
↓
Autolink matching mentions
↓
Validate every link
↓
Restore false positives from backup
⚠ Challenges
Four notes came back damaged in a subtle way: raw terminal and Vim bracket syntax in lab session logs and cheat sheets looked enough like wiki-link syntax that the linker mangled it. Validation caught it, but it only caught it because I checked link balance file by file instead of trusting the summary counts.
📚 What I Learned
I learned that “0 missing targets” in a validation report is not the same as “nothing went wrong.” The real risk in bulk text automation isn’t missing links, it’s confidently correct-looking output that quietly broke something the validator wasn’t checking for.
➡ Next Steps
- Add personal lab examples to high-value terms like SSH, SUID, IDOR, CSRF, and lateral movement
- Consider additional topic maps: Identity and Access, Malware and Adversary Behavior, Detection Engineering, Cloud and Infrastructure
- Keep the pre-link backup archive until I’m confident no other false positives exist
🧠 Reflection
This was a good reminder that automation validated against the wrong criteria is still unvalidated. I was checking “did links break,” not “did the linker rewrite something it should have left alone” — and both matter.
🧩 Lessons Learned
What worked
Taking a full backup before running a bulk change across 485 files, and skipping frontmatter, code blocks, URLs, and existing links by rule.
What broke
Four notes with raw terminal/Vim bracket syntax got misread as wiki-link syntax and needed manual restoration.
Why it broke
Text that looks like [[...]] syntax isn’t always a wiki link — sometimes it’s just what a terminal session or a Vim cheat sheet looks like.
Fix / takeaway
Validate for “unexpected changes,” not just “expected links present” — a clean link-count report can still hide corrupted content.
📈 Skill Progression Context
This supports my cybersecurity progression because running a large automated change safely — backup first, narrow rules, validate for the failure modes you didn’t anticipate — is the same operational discipline that separates a careful defender from one who trusts their own tooling too quickly.
😄 TL;DR
348 terms, 3,820 new links, and four notes that had to be rescued from a false-positive linker — the backup is what saved it.
