🔄 Topic

Letting Hermes generate real financial-adjacent PDFs — an Amazon purchase register, a shareable spending note — meant building a supervision layer that refuses to let “the runtime found the skill” count as “the skill actually ran,” and refuses to let a scoped data source quietly expand into a claim it doesn’t support.


🎯 Goal

Get Hermes producing real, evidence-backed PDF reports from a single named data source, with a fail-closed gate that only trusts a completed, traced execution — never a plausible-looking absence of errors — and with the report’s own scope stated honestly rather than inflated.


🛠 What I Did

I built the evidence-gated PDF pipeline in stages, and each stage closed a specific way the system could have quietly claimed more than it had actually verified.

Main areas covered:

  • shipped hermes-pdf-supervision v1.0.0: a sequential purchase-report pipeline plus a consent-first generator for a separate personal document type, where dependent PDF stages run in one foreground pipeline and Hermes owns final execution and gating even after invoking a checked generator to repair inline-code syntax
  • generated and validated three real PDFs from a single, explicitly scoped source: a private Amazon card capture only — no Mail data, no bank data folded in, even though both would have been easy to blend in for a “more complete” picture
  • every generated PDF passed an automated gate (pdf_gate.py --render-all --min-footer-y 30) plus an independent page-count check, and covers, dense content pages, and final pages were rendered and visually inspected rather than trusting the automated gate alone
  • explicitly labeled what the output is and is not: an Amazon order-card package for a specific one-year window, not a complete spend ledger — with mail-receipt discovery and bank-statement reconciliation named as separate, still-unscoped future stages rather than implied as already covered
  • followed up with hermes-pdf-runtime-proof (v0.7.1 of the supervision skill), which structurally separates three distinct claims that had previously been at risk of blurring together: skill discovery (the runtime can find and read the skill), behavioral proof (a completed safe canary with a visible execution trace), and a no-response process outcome (an inconclusive liveness signal that is never treated as proof)
  • bounded the verification canary itself so a hung test can’t consume workstation resources indefinitely, permitting termination only of the specific verified canary process
  • ran the actual canary and got an inconclusive result — no response within 90 seconds — and rather than treating “the runtime can see the skill” as good enough to call it verified, documented plainly that a completed behavioral trace was not yet available and that no one should claim the model executed the skill until a later, successful canary actually produces one

🔗 Key Cybersecurity Connections

This is the false-assurance pattern from earlier this month, now built directly into the architecture of a new capability instead of being found after the fact: “the system can locate this skill” and “the system actually executed this skill and left proof” are two different claims, and collapsing them is exactly how false confidence gets baked into a fail-open system. Structurally separating discovery, proof, and inconclusive-non-proof means the gate can’t be satisfied by a technically-true-but-insufficient signal.

Scoping the data source explicitly — Amazon card data only, mail and bank data named as future, separate work — is the same discipline applied to a report’s claimed completeness rather than to code execution: a report that quietly draws from more sources than it discloses, or is presented as more comprehensive than it is, is its own kind of false assurance, just aimed at a human reader instead of at the calling system. Both problems get the same fix: say exactly what was verified, and say exactly what wasn’t, rather than letting either side round up.


🔍 Investigation Questions

  • Does “the system can find this capability” ever get treated as equivalent to “the system successfully used this capability”?
  • Is a no-response or inconclusive verification result ever silently counted as a pass?
  • Does a generated report state precisely which data sources fed it, or does it imply broader coverage than what was actually included?
  • Can a hung verification canary consume unbounded resources, or is its termination scoped to exactly the verified test process?
  • Was output actually inspected (rendered, viewed) or only checked by an automated gate that could itself have blind spots?

🚨 Detection Opportunities

Checks for an evidence-gated generation pipeline:

  • skill/capability discovery treated as proof of successful execution
  • a no-response or timeout outcome from a verification canary silently logged as a pass
  • a generated report’s stated scope broader than its actual data sources
  • an unbounded verification process risking runaway resource consumption
  • automated output validation with no corresponding manual/visual spot-check

Example:

project=hermes-pdf-supervision
signal=skill_discovery_treated_as_execution_proof
risk_area=false_assurance_in_evidence_gate_design
triage=separate_discovery_proof_and_inconclusive_states_explicitly

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is fail-closed evidence-gate design and data-scope honesty for an automated generation pipeline, not an adversary technique.


🗺 Visual Investigation Diagram

Hermes needs to generate real financial-adjacent PDFs
    ↓
v1.0.0: sequential pipeline, single scoped data source (Amazon only)
    ↓
Automated gate + independent page check + manual visual inspection
    ↓
Report explicitly labeled: NOT a complete spend ledger
    ↓
v0.7.1: separate discovery vs behavioral proof vs inconclusive
    ↓
Canary run: no response in 90s — inconclusive, not success
    ↓
Documented honestly: do not claim execution until proof exists

⚠ Challenges

The harder discipline was writing down the inconclusive canary result as inconclusive, rather than letting “the skill is discoverable and the always-loaded instruction requires it” read as good enough in the handoff. It would have been easy to imply the gate was fully proven when only half of it — discovery — actually was.


📚 What I Learned

I learned that an evidence gate’s own design can quietly reintroduce the false-assurance problem it’s meant to prevent, if “can find it” and “proved it ran” aren’t kept as structurally distinct claims. I also learned that scoping a generated report’s data sources explicitly, and saying plainly what it doesn’t yet cover, is the same honesty discipline as a code-level evidence gate — just applied to what a human reader might otherwise assume.


➡ Next Steps

  • Run a successful canary to actually close the behavioral-proof gap left open by the inconclusive 90-second run
  • Scope mail-receipt discovery and bank-statement reconciliation as their own explicitly-gated stages when built, not folded silently into the existing pipeline
  • Apply the discovery/proof/inconclusive three-way split to other skill-verification gates in the stack
  • Periodically re-inspect generated PDFs visually, not just via the automated gate, as the pipeline evolves

🧠 Reflection

The inconclusive canary result was, in a small way, the more important finding of the day — not because anything broke, but because writing “this is not yet proof” honestly, right when it would have been easiest to round up to “verified,” is exactly the habit that keeps an evidence gate meaning something.


🧩 Lessons Learned

What worked

Structurally separating skill discovery from behavioral proof, scoping generated reports to their actual data sources, and documenting an inconclusive verification result as inconclusive rather than rounding it up.

What broke

Nothing shipped broken — the inconclusive canary was caught and reported honestly rather than treated as a passing verification.

Why it mattered

An evidence gate that conflates “discoverable” with “proven executed” reintroduces the exact false-assurance risk it exists to prevent, and a report that implies broader coverage than its actual sources misleads its reader the same way.

Fix / takeaway

Keep discovery, proof, and inconclusive outcomes as three distinct, honestly-labeled states, and scope every generated report to exactly the data sources it actually used.


📈 Skill Progression Context

This supports my cybersecurity progression because fail-closed evidence-gate architecture, honest scope disclosure, and resisting the urge to round an inconclusive result up to a pass are core integrity practices for any automated system whose output people will actually rely on.


😄 TL;DR

Hermes can generate real purchase-report PDFs now — but the gate that verifies it actually ran the skill won’t accept “the runtime found it” as proof, and an inconclusive test stayed labeled inconclusive instead of quietly becoming a pass.