π Day 160 β Tailscale Broke My Stack: Node Identity, Hostnames, and Private Services
π Topic
My private network layer stopped cooperating: connections dropping daily, an iOS companion endpoint that vanished, and services that worked yesterday returning nothing today. The root causes were about node identity and hostnames, not about the apps.
π― Goal
Understand why the tailnet kept βlosingβ my machine, fix the private access to every service, and end with one stable node identity everything agrees on.
π What I Did
I debugged the overlay network like an incident.
Main areas covered:
- investigated why remote connections needed re-setup daily and whether the MacBookβs IP changing was to blame
- discovered the real issue was node identity: multiple Tailscale installations meant services were bound to a node that was not the one currently online
- migrated the whole stack to the GUI appβs node (macbook-pro) and reconciled hostnames everywhere they were referenced
- repaired private access to Hermes and n8n through the tailnet
- restored Open WebUI so the phone could reach the local models again
- traced the dead iOS companion URL to the same stale-hostname problem and updated the remote hub URL
- documented the reconciliation so the next hostname change is a checklist, not an outage
π Key Cybersecurity Connections
Private overlay networks are a security control β and like any control, they fail as availability first. Every one of these outages was an identity problem: which node is βmy laptop,β which hostname do clients trust, which certificate matches. That is the same class of problem as certificate pinning breaks and DNS drift in any enterprise.
A subtle risk: when private access breaks repeatedly, the temptation grows to expose the service publicly βjust for now.β Fixing reliability is what protects the security decision.
π Investigation Questions
- Which node identity is each service actually bound to?
- Do all clients reference the same hostname, or a mix of stale ones?
- Is the daily disconnect a network problem or an identity/key problem?
- What breaks when the machine sleeps or the app restarts?
- Where are tailnet hostnames hardcoded across my projects?
π¨ Detection Opportunities
Checks for a private service mesh:
- service bound to a node that is offline while a twin node is online
- clients resolving different hostnames for the same service
- daily reconnect patterns suggesting identity churn rather than link failure
- private endpoint returning nothing while the service process is healthy
- hostname references diverging across configs and docs
Example:
project=tailnet-stack
signal=service_bound_to_offline_node_identity
risk_area=private_access_availability
triage=enumerate_nodes_reconcile_hostnames_rebind_services
π§ MITRE ATT&CK Techniques
No direct mapping claimed. This is availability and identity hygiene for a private network control.
πΊ Visual Investigation Diagram
"It stopped working again"
β
Enumerate tailnet nodes
β
Find duplicate node identities
β
Pick one canonical node
β
Rebind services + reconcile hostnames
β
Document the checklist
β
Stable private access
β Challenges
The misleading part was the symptom: it looked like an IP problem, and tailnets are designed so IPs do not matter. Letting go of the wrong hypothesis and enumerating node identities instead was the turning point.
π What I Learned
I learned that in an overlay network, identity is the address. Two installs of the same client means two machines as far as the tailnet cares, and services do not follow you between them.
β‘ Next Steps
- Keep exactly one Tailscale installation per machine
- Grep projects for tailnet hostnames after any node change
- Add a quick reachability check for each private service
- Resist any βexpose it publicly for nowβ shortcut
π§ Reflection
Network debugging felt like the purest analyst work of the week: symptoms, wrong hypothesis, evidence, real cause, remediation, and a checklist so it never costs a day again.
π§© Lessons Learned
What worked
Treating the outage as an investigation with hypotheses instead of restarting things at random.
What broke
Services bound to a stale node identity from a second Tailscale install.
Why it broke
Two installations silently created two node identities for one laptop.
Fix / takeaway
One node identity per machine, one canonical hostname, and a reconciliation checklist for changes.
π Skill Progression Context
This supports my cybersecurity progression because identity-versus-address thinking, DNS drift, and availability of security controls are everyday realities of defending real networks.
π TL;DR
The network was fine; my laptop had an identity crisis.
