πŸ”„ Topic

Finding that an unattended document-reconciliation job was repeatedly choosing the same high-priority files and silently starving the rest of its registry.


🎯 Goal

Make a nightly maintenance routine respond to genuine urgency without letting always-important documents monopolize every turn.


πŸ›  What I Did

Production evidence across six unattended nights showed a clear pattern: canonical status documents became β€œbehind today” every day, so a priority-first sorter selected them repeatedly. Other registered documents never received a reconciliation turn.

I changed the ranking order to:

  1. genuine staleness first;
  2. longest time since a document last received a turn, with never-reconciled items first;
  3. document priority;
  4. modification time.

The proof was intentionally adversarial. A new regression test fails under the old ordering with the message that a canonical document won every night, and passes under the new ordering after a lower-priority document receives a turn in a simulated multi-night sequence.


πŸ”— Key Cybersecurity Connections

Security maintenance is only useful if coverage is real. Starvation can leave inventories, runbooks, or recovery notes stale indefinitely while a system appears busy and healthy.


πŸ” Investigation Questions

  • Which items receive maintenance turns over time?
  • Is priority masking starvation?
  • Can a lower-priority but never-reviewed item ever win?
  • Does urgent work still preempt the fairness rule?
  • Can the old faulty sort order make the test fail?

🚨 Detection Opportunities

Track:

  • documents with no reconciliation turn within a defined period
  • repeated selection of the same target
  • registry entries that are never selected
  • stale items that were deferred despite higher urgency

Example:

event=reconcile_starvation
document=knowledge/project-index.md
turns_since_last=never
action=raise_priority_or_rebalance_selector

🧭 MITRE ATT&CK Techniques

This is operational resilience rather than a direct ATT&CK mapping. It supports accurate defensive documentation and maintenance evidence.


πŸ—Ί Visual Investigation Diagram

Old selector
    ↓
Canonical document behind today
    ↓
Same document wins again
    ↓
Other registry entries starve

New selector
    ↓
Urgency β†’ time since last turn β†’ priority β†’ age
    ↓
Coverage becomes observable

⚠ Challenges

Priority feels like safety, but a static priority order can be unfair forever. The fix had to preserve true urgency while correcting ties that were happening every night.


πŸ“š What I Learned

I learned that fairness is measurable. If a recurring job chooses work, its selection history is part of its security and reliability evidence.


➑ Next Steps

  • Keep the starvation regression test
  • Review registry coverage periodically
  • Treat β€œnever selected” as a first-class maintenance signal

🧠 Reflection

A system can be busy every night and still fail to care for part of what it owns.


🧩 Lessons Learned

What worked

Ranking staleness before fairness and fairness before fixed priority.

What broke

Assuming a priority-first sorter would eventually rotate naturally.

Why it broke

The same canonical documents regenerated their urgency every day.

Fix / takeaway

Test selection over time, not only the result of one run.


πŸ“ˆ Skill Progression Context

This connects scheduling logic to defensive operations: coverage, evidence, and predictable recovery all need to be designed, not assumed.


πŸ˜„ TL;DR

When an unattended job always picks the same β€œimportant” work, fairness becomes a reliability controlβ€”not a nice-to-have.