πŸ”„ Topic

My private network layer stopped cooperating: connections dropping daily, an iOS companion endpoint that vanished, and services that worked yesterday returning nothing today. The root causes were about node identity and hostnames, not about the apps.


🎯 Goal

Understand why the tailnet kept β€œlosing” my machine, fix the private access to every service, and end with one stable node identity everything agrees on.


πŸ›  What I Did

I debugged the overlay network like an incident.

Main areas covered:

  • investigated why remote connections needed re-setup daily and whether the MacBook’s IP changing was to blame
  • discovered the real issue was node identity: multiple Tailscale installations meant services were bound to a node that was not the one currently online
  • migrated the whole stack to the GUI app’s node (macbook-pro) and reconciled hostnames everywhere they were referenced
  • repaired private access to Hermes and n8n through the tailnet
  • restored Open WebUI so the phone could reach the local models again
  • traced the dead iOS companion URL to the same stale-hostname problem and updated the remote hub URL
  • documented the reconciliation so the next hostname change is a checklist, not an outage

πŸ”— Key Cybersecurity Connections

Private overlay networks are a security control β€” and like any control, they fail as availability first. Every one of these outages was an identity problem: which node is β€œmy laptop,” which hostname do clients trust, which certificate matches. That is the same class of problem as certificate pinning breaks and DNS drift in any enterprise.

A subtle risk: when private access breaks repeatedly, the temptation grows to expose the service publicly β€œjust for now.” Fixing reliability is what protects the security decision.


πŸ” Investigation Questions

  • Which node identity is each service actually bound to?
  • Do all clients reference the same hostname, or a mix of stale ones?
  • Is the daily disconnect a network problem or an identity/key problem?
  • What breaks when the machine sleeps or the app restarts?
  • Where are tailnet hostnames hardcoded across my projects?

🚨 Detection Opportunities

Checks for a private service mesh:

  • service bound to a node that is offline while a twin node is online
  • clients resolving different hostnames for the same service
  • daily reconnect patterns suggesting identity churn rather than link failure
  • private endpoint returning nothing while the service process is healthy
  • hostname references diverging across configs and docs

Example:

project=tailnet-stack
signal=service_bound_to_offline_node_identity
risk_area=private_access_availability
triage=enumerate_nodes_reconcile_hostnames_rebind_services

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is availability and identity hygiene for a private network control.


πŸ—Ί Visual Investigation Diagram

"It stopped working again"
    ↓
Enumerate tailnet nodes
    ↓
Find duplicate node identities
    ↓
Pick one canonical node
    ↓
Rebind services + reconcile hostnames
    ↓
Document the checklist
    ↓
Stable private access

⚠ Challenges

The misleading part was the symptom: it looked like an IP problem, and tailnets are designed so IPs do not matter. Letting go of the wrong hypothesis and enumerating node identities instead was the turning point.


πŸ“š What I Learned

I learned that in an overlay network, identity is the address. Two installs of the same client means two machines as far as the tailnet cares, and services do not follow you between them.


➑ Next Steps

  • Keep exactly one Tailscale installation per machine
  • Grep projects for tailnet hostnames after any node change
  • Add a quick reachability check for each private service
  • Resist any β€œexpose it publicly for now” shortcut

🧠 Reflection

Network debugging felt like the purest analyst work of the week: symptoms, wrong hypothesis, evidence, real cause, remediation, and a checklist so it never costs a day again.


🧩 Lessons Learned

What worked

Treating the outage as an investigation with hypotheses instead of restarting things at random.

What broke

Services bound to a stale node identity from a second Tailscale install.

Why it broke

Two installations silently created two node identities for one laptop.

Fix / takeaway

One node identity per machine, one canonical hostname, and a reconciliation checklist for changes.


πŸ“ˆ Skill Progression Context

This supports my cybersecurity progression because identity-versus-address thinking, DNS drift, and availability of security controls are everyday realities of defending real networks.


πŸ˜„ TL;DR

The network was fine; my laptop had an identity crisis.