πŸ”„ Topic

Building encrypted off-host backup for Hermes state, then proving the restore on the disaster machine without interrupting the live messaging bridge.


🎯 Goal

Protect the irreplaceable local state behind Hermes while proving that a restore is structurally sound and safe to inspect.


πŸ›  What I Did

I added encrypted off-host backup using a restic repository reachable over Tailscale, along with a restore-verification path that runs on the separate recovery machine.

The design solved several real problems:

  • the live SQLite database uses WAL mode, so copying the file directly produced a demonstrably inconsistent snapshot;
  • the backup uses a consistent SQLite snapshot rather than the live file;
  • the verifier reads restored tables and checks structure without starting a second messaging bridge;
  • derived FTS5-index complaints are treated separately from non-derived database corruption;
  • the restore was tested on the recovery machine, not merely on the laptop that created it.

The verified restore read 2,606 files, completed 14 checks, and inspected 21 tables with 64,977 rows. The live bridge process and credential-file hash were unchanged before and after the test.


πŸ”— Key Cybersecurity Connections

This is availability and recovery engineering. Encryption protects the backup’s confidentiality; structural restore verification protects against the more dangerous illusion of having a backup that cannot actually recover the system.


πŸ” Investigation Questions

  • Is the backup consistent with the live database’s storage mode?
  • Can the recovery machine restore it independently?
  • Does verification accidentally start a second client with live credentials?
  • Are integrity warnings derived-index noise or evidence of real corruption?
  • Did the test read the restored data, not only list an archive?

🚨 Detection Opportunities

Track:

  • snapshot failure or missing backup contents
  • restore verification failures by class
  • unexpected bridge PID or credential hash changes during verification
  • retention failures and backup-size growth

Example:

event=restore_verification_failed
class=non_derived_sqlite_integrity_error
action=halt_recovery_claim_and_investigate_snapshot

🧭 MITRE ATT&CK Techniques

Relevant defensive context:

  • T1486 β€” Data Encrypted for Impact, because recoverable encrypted backups reduce ransomware impact
  • T1490 β€” Inhibit System Recovery, because independent restore proof protects against assumptions about recovery readiness

πŸ—Ί Visual Investigation Diagram

Live Hermes state
    ↓
Consistent SQLite snapshot
    ↓
Encrypted off-host restic backup
    ↓
Restore on separate recovery machine
    ↓
Structural checks, no live bridge start

⚠ Challenges

A backup command can complete while still preserving a torn database or a verifier can reject healthy data for the wrong compatibility reason. Both failure modes needed to be measured rather than guessed.


πŸ“š What I Learned

I learned that backup verification must be specific about what it proves. A file count, an archive listing, and a restored, readable application state are different levels of evidence.


➑ Next Steps

  • Keep retention bounded and observable
  • Re-run restoration after important state-schema changes
  • Preserve the separate recovery environment as part of the design

🧠 Reflection

The safest recovery test is one that proves the backup works without touching the live service it is meant to protect.


🧩 Lessons Learned

What worked

Snapshotting SQLite correctly, encrypting off-host storage, and restoring elsewhere.

What broke

Trusting raw live-database copies and simplistic pipeline checks.

Why it broke

Live data has transactional and compatibility details that a file copy does not capture.

Fix / takeaway

Back up consistent state and prove recovery from a separate machine.


πŸ“ˆ Skill Progression Context

This gave me direct experience with recovery planning, encryption, SQLite consistency, and evidence-based disaster-readiness checks.


πŸ˜„ TL;DR

An encrypted backup became trustworthy only after it restored cleanly on another machine without disturbing the live bridge.