đ Day 147 â 'Build Succeeded' Isn't Proof: Writing a Real UI Test for the Auth Flow
đ Topic
Following up on the pairing-token auth I added to my agent daemon and iOS app: instead of leaving it at âcompiles successfully,â I wrote a real end-to-end UI test that drives the actual app against the actual live daemon.
đŻ Goal
Prove the authenticated flow works the way a real user experiences it, not just the way the compiler thinks it should.
đ What I Did
I built a genuine XCUITest that launches the app, navigates the real UI, enters the pairing token, and asserts the connection state changes â against the live daemon over the real network path, not a mock.
Main areas covered:
- adding a UI test target to a project that had none before
- writing a test that launches the app, handles the tab barâs overflow menu, types the token into a secure field, and taps the real âRefresh Diagnosticsâ button
- passing the test token in through an environment variable at the scheme level, never hardcoded in source
- adding accessibility identifiers so the test targets real UI elements instead of guessing at auto-generated labels
- fixing two real failures the test surfaced, neither of which was a bug in the app
- keeping the test in the repo permanently as a standing regression check
đ Key Cybersecurity Connections
A build succeeding tells you the code compiles. It tells you nothing about whether the authentication flow a real user depends on actually completes correctly end to end. This is the same gap that shows up in security testing generally: a control existing in code is not the same claim as a control working when exercised the way an attacker or a legitimate user actually would exercise it. I treated âauth worksâ as a claim that needed a real test, not an inference from a successful compile.
đ Investigation Questions
- Does âbuild succeededâ actually verify the behavior I care about, or just that the syntax is valid?
- Is the test exercising the real network path, or a mock that could hide a real integration failure?
- Are test credentials injected safely (environment variable) or accidentally committed to source?
- When the test fails, is it because the app is broken, or because the testâs assumptions about UI structure were wrong?
đš Detection Opportunities
Potential monitoring ideas â applied to my own test suite instead of a live system:
- security-relevant flows (auth, permissions, credential handling) with zero test coverage
- test credentials hardcoded in source instead of injected at runtime
- a passing test suite that never actually exercises the real backend
- UI tests that silently rot because they were never converted from a one-off script into a standing test
Example:
project=agent-companion-ios
change_type=security_flow_test_added
risk_area=untested_authentication_path
triage=confirm_test_exercises_real_backend_not_a_mock
đ§ MITRE ATT&CK Techniques
Not directly applicable â this is verification engineering, not adversary behavior. No mappings claimed here.
đș Visual Investigation Diagram
Build succeeds (compiles)
â
Assumption: auth flow works
â
Write real XCUITest against live daemon
â
Two failures surfaced (tab overflow, label mismatch)
â
Fix test assumptions, not the app
â
Third run: real end-to-end success, kept as regression test
â Challenges
Both failures I hit initially looked like app bugs and werenât. The tab bar folded a sixth tab under âMore,â and SwiftUIâs LabeledContent merged the label into the accessibility string so a strict equality check failed even though the state was correct. It would have been easy to âfixâ the app to match a wrong assumption in the test instead of fixing the test.
đ What I Learned
I learned to be suspicious of my own first failure diagnosis. Both times, the instinct to change the app was wrong â the testâs assumptions about UI structure were the actual bug. Verifying which side is actually broken matters more than fixing the first thing that looks broken.
⥠Next Steps
- Consider capturing the simulator-driving pattern (tab-bar folding, accessibility identifiers, env-var token injection) as a reusable pattern if more UI tests get written later
- Keep this test as a standing regression check whenever the auth flow changes
- Extend similar real-device verification to the other âbuild-verified onlyâ features from the same session
đ§ Reflection
This was a good reminder that verification has a hierarchy: compiles, then runs, then behaves correctly under real conditions a user will actually hit. Skipping straight from the first to the third is how confidently wrong software ships.
đ§© Lessons Learned
What worked
Writing a real UI test against the live backend instead of accepting âit buildsâ as sufficient proof for a security-relevant flow.
What broke
Two of my own testâs assumptions about UI structure â not the app itself.
Why it broke
Real UI structure (tab overflow, merged accessibility labels) doesnât always match what youâd guess from reading the view code.
Fix / takeaway
When a test fails, donât assume the app is wrong â verify which side actually is before changing anything.
đ Skill Progression Context
This supports my cybersecurity progression because distinguishing âcompilesâ from âverified workingâ is the same discipline needed when validating that a security control actually functions under real conditions, not just in theory.
đ TL;DR
âBuild succeededâ isnât proof anything works â wrote a real end-to-end test for the auth flow and found two wrong assumptions in my own test before it passed for real.
