🔄 Topic

A “verified Apple Mail delivery” feature was reporting success the moment a send call returned — not when the message actually arrived. Fixing that, and finding secrets leaking into diagnostic evidence along the way, closed both gaps in the same session.


🎯 Goal

Make “delivered” mean the message is actually confirmed present, and make sure the evidence proving that never leaks anything it shouldn’t.


🛠 What I Did

I treated my own delivery-reporting feature as a claim that needed proof, not a status that needed logging.

Main areas covered:

  • found that Mail verification declared success based on the send call returning without error — exactly the false-assurance pattern from the supervisor audit weeks earlier, now in a different subsystem
  • made the verification evidence-gated: success now requires confirming the message’s actual presence, not just an unhandled-exception-free send
  • built the verified-delivery reporting path so a completed report includes the evidence that it landed, not just an assertion that it was asked to be sent
  • while building the evidence path, found CLI-flag secrets leaking into route-doctor listener evidence output and redacted them before they could appear in any diagnostic artifact
  • documented the “send is not delivery” lesson explicitly, so future features default to evidence-gating rather than needing the same bug found twice
  • tightened the artifact-link HTTP clock in a related test to be deterministic, removing a source of flaky evidence in the delivery test suite

🔗 Key Cybersecurity Connections

“The function call returned successfully” and “the real-world effect happened” are different claims, and conflating them is the same false-assurance failure mode I hardened the task supervisor against — just showing up again in a new place, because the lesson had not yet generalized past the system where it was first found.

The secret-redaction fix is a straightforward but easy-to-miss one: diagnostic and evidence-gathering code paths are exactly where secrets leak, because they are built to capture everything, and “everything” usually includes things that should never be written down.


🔍 Investigation Questions

  • Does “delivered” mean confirmed-present, or just send-attempted-without-error?
  • What evidence actually proves delivery, and is it captured?
  • Do diagnostic or evidence-gathering paths ever capture secrets meant only for authentication?
  • Is this false-assurance pattern present anywhere else it has not yet been looked for?
  • Are delivery tests deterministic, or dependent on real-world timing?

🚨 Detection Opportunities

Checks for delivery and evidence pipelines:

  • success reported on send-without-exception rather than confirmed presence
  • diagnostic evidence containing CLI flags, tokens, or other secret-shaped strings
  • delivery tests flaking due to non-deterministic clocks or timers
  • a “verified” claim anywhere with no corresponding evidence artifact
  • the same false-assurance pattern recurring in a subsystem it was already fixed in elsewhere

Example:

project=hermes-mail-delivery
signal=delivery_reported_on_send_success_only
risk_area=false_assurance
triage=require_confirmed_presence_evidence_before_success

🧭 MITRE ATT&CK Techniques

Possible mapping for the secret-leak finding:

  • T1552 — Unsecured Credentials

No direct mapping claimed for the false-assurance fix; it is a reliability and integrity issue rather than an adversary technique.


🗺 Visual Investigation Diagram

Send call returns
    ↓
Old: report success here
    ↓
New: confirm actual presence
    ↓
Evidence attached to the report
    ↓
Diagnostic path audited for secrets
    ↓
CLI-flag secrets redacted before capture

⚠ Challenges

The uncomfortable recognition was that this was the same bug class already fixed once, in the supervisor, weeks ago — and it had not yet occurred to me to go looking for it elsewhere. A lesson learned in one subsystem does not automatically protect the next one.


📚 What I Learned

I learned that false-assurance patterns need a checklist, not a memory. “Success reported without proof of real-world effect” should be a standing question asked of every feature, not a bug fixed once and trusted to stay fixed everywhere.


➡ Next Steps

  • Write down the false-assurance pattern as a standing review question for new features
  • Audit other notification and delivery paths for the same send-versus-delivered gap
  • Add a secret-scanning pass over diagnostic and evidence-capture code specifically
  • Keep deterministic clocks as the default in any test touching timing

🧠 Reflection

Finding the same category of bug twice, in two unrelated systems, was more useful than finding it once — it turned a one-off fix into a pattern I now actively hunt for.


🧩 Lessons Learned

What worked

Treating “delivered” as a claim requiring evidence, and auditing the evidence path itself for leaks.

What broke

Delivery success reported on send-without-error, and secrets leaking into diagnostic evidence.

Why it broke

A pattern already fixed elsewhere had not yet been generalized into a standing check for new features.

Fix / takeaway

Evidence-gate every delivery claim, audit diagnostic paths for secrets by default, and turn one-off fixes into standing review questions.


📈 Skill Progression Context

This supports my cybersecurity progression because distinguishing attempted action from confirmed effect, and auditing diagnostic paths for credential leakage, are both everyday concerns in reliable, secure system design.


😄 TL;DR

The mail said “sent”; now it has to prove “arrived” — and its evidence trail no longer leaks secrets.