🔄 Topic

Following up on the pairing-token auth I added to my agent daemon and iOS app: instead of leaving it at “compiles successfully,” I wrote a real end-to-end UI test that drives the actual app against the actual live daemon.


🎯 Goal

Prove the authenticated flow works the way a real user experiences it, not just the way the compiler thinks it should.


🛠 What I Did

I built a genuine XCUITest that launches the app, navigates the real UI, enters the pairing token, and asserts the connection state changes — against the live daemon over the real network path, not a mock.

Main areas covered:

  • adding a UI test target to a project that had none before
  • writing a test that launches the app, handles the tab bar’s overflow menu, types the token into a secure field, and taps the real “Refresh Diagnostics” button
  • passing the test token in through an environment variable at the scheme level, never hardcoded in source
  • adding accessibility identifiers so the test targets real UI elements instead of guessing at auto-generated labels
  • fixing two real failures the test surfaced, neither of which was a bug in the app
  • keeping the test in the repo permanently as a standing regression check

🔗 Key Cybersecurity Connections

A build succeeding tells you the code compiles. It tells you nothing about whether the authentication flow a real user depends on actually completes correctly end to end. This is the same gap that shows up in security testing generally: a control existing in code is not the same claim as a control working when exercised the way an attacker or a legitimate user actually would exercise it. I treated “auth works” as a claim that needed a real test, not an inference from a successful compile.


🔍 Investigation Questions

  • Does “build succeeded” actually verify the behavior I care about, or just that the syntax is valid?
  • Is the test exercising the real network path, or a mock that could hide a real integration failure?
  • Are test credentials injected safely (environment variable) or accidentally committed to source?
  • When the test fails, is it because the app is broken, or because the test’s assumptions about UI structure were wrong?

🚹 Detection Opportunities

Potential monitoring ideas — applied to my own test suite instead of a live system:

  • security-relevant flows (auth, permissions, credential handling) with zero test coverage
  • test credentials hardcoded in source instead of injected at runtime
  • a passing test suite that never actually exercises the real backend
  • UI tests that silently rot because they were never converted from a one-off script into a standing test

Example:

project=agent-companion-ios
change_type=security_flow_test_added
risk_area=untested_authentication_path
triage=confirm_test_exercises_real_backend_not_a_mock

🧭 MITRE ATT&CK Techniques

Not directly applicable — this is verification engineering, not adversary behavior. No mappings claimed here.


đŸ—ș Visual Investigation Diagram

Build succeeds (compiles)
    ↓
Assumption: auth flow works
    ↓
Write real XCUITest against live daemon
    ↓
Two failures surfaced (tab overflow, label mismatch)
    ↓
Fix test assumptions, not the app
    ↓
Third run: real end-to-end success, kept as regression test

⚠ Challenges

Both failures I hit initially looked like app bugs and weren’t. The tab bar folded a sixth tab under “More,” and SwiftUI’s LabeledContent merged the label into the accessibility string so a strict equality check failed even though the state was correct. It would have been easy to “fix” the app to match a wrong assumption in the test instead of fixing the test.


📚 What I Learned

I learned to be suspicious of my own first failure diagnosis. Both times, the instinct to change the app was wrong — the test’s assumptions about UI structure were the actual bug. Verifying which side is actually broken matters more than fixing the first thing that looks broken.


➡ Next Steps

  • Consider capturing the simulator-driving pattern (tab-bar folding, accessibility identifiers, env-var token injection) as a reusable pattern if more UI tests get written later
  • Keep this test as a standing regression check whenever the auth flow changes
  • Extend similar real-device verification to the other “build-verified only” features from the same session

🧠 Reflection

This was a good reminder that verification has a hierarchy: compiles, then runs, then behaves correctly under real conditions a user will actually hit. Skipping straight from the first to the third is how confidently wrong software ships.


đŸ§© Lessons Learned

What worked

Writing a real UI test against the live backend instead of accepting “it builds” as sufficient proof for a security-relevant flow.

What broke

Two of my own test’s assumptions about UI structure — not the app itself.

Why it broke

Real UI structure (tab overflow, merged accessibility labels) doesn’t always match what you’d guess from reading the view code.

Fix / takeaway

When a test fails, don’t assume the app is wrong — verify which side actually is before changing anything.


📈 Skill Progression Context

This supports my cybersecurity progression because distinguishing “compiles” from “verified working” is the same discipline needed when validating that a security control actually functions under real conditions, not just in theory.


😄 TL;DR

“Build succeeded” isn’t proof anything works — wrote a real end-to-end test for the auth flow and found two wrong assumptions in my own test before it passed for real.