🔄 Topic

A real audit of the phone-approval gate from a few weeks back found several concrete weaknesses: an approval queue keyed on the wrong identity, a policy hook that could be knocked over silently, and a webhook sitting on the network where nothing needed it to be.


🎯 Goal

Close the specific gaps a hostile actor — or just bad luck — could use against the approval mechanism that is supposed to be the last word before anything external happens.


🛠 What I Did

I went through the approval queue’s actual failure modes one at a time.

Main areas covered:

  • keyed the approval queue on the sender, not the message — the previous keying could let approval of one message get confused with a different message from the same conversation
  • minted callback tokens for approvals instead of relying on freeform replies, and auto-denied automated Italian mail that had no business reaching the approval flow at all
  • moved approval delivery to Telegram polling instead of a reachable URL — a publicly reachable webhook for approving privileged actions is attack surface that does not need to exist
  • authenticated the sender-policy hook and started backing up the policy file, so a bad actor can’t spoof a policy-hook call and a bad edit can’t destroy the policy with no way back
  • made sure a policy-hook outage never turns into a lost decision — if the hook is unreachable, the system waits, it doesn’t silently drop the approval request
  • bounded and expired the approval queue so requests don’t linger forever, while making sure expiry reads as “this specific request timed out,” not as a blanket block
  • added a status surface reporting the sender-approval backlog, and treated a dead approval poller itself as an outage worth alerting on
  • wrote a dedicated test covering the sender-policy hook specifically, and logged four distinct findings from this repair as standing lessons

🔗 Key Cybersecurity Connections

Keying an approval by message instead of sender is a subtle identity bug with a real consequence: two messages in flight could get their approvals crossed. That is the same category of mistake as session-fixation or token-confusion bugs — the system asked “was a thing approved” instead of “was this specific thing, from this specific principal, approved.”

Removing the reachable webhook in favor of polling is attack-surface reduction in its purest form: nothing can hit an endpoint that does not exist. And treating “the approval poller died” as an outage closes the same gap the Google-auth watchdog closed yesterday — a control that can silently stop working is not a control you can rely on.


🔍 Investigation Questions

  • Does the approval queue key on the requester’s identity, or something looser?
  • Can a policy-hook outage cause a decision to be silently lost or silently allowed?
  • Is any approval-relevant endpoint reachable from outside that doesn’t need to be?
  • Is the policy file backed up, and is the hook that reads it authenticated?
  • Does the system alert when its own approval mechanism stops functioning?

🚨 Detection Opportunities

Checks for an approval-gated pipeline:

  • an approval resolved for a request other than the one it was issued for
  • a policy-hook call with no authentication credential
  • an approval-relevant endpoint reachable without an active session or token
  • the approval poller silent past its expected heartbeat
  • a queued approval expiring silently with no distinguishable “timed out” signal

Example:

project=approval-queue-hardening
signal=approval_resolved_for_wrong_request
risk_area=identity_confusion_in_authorization_flow
triage=verify_queue_key_is_sender_and_request_scoped

🧭 MITRE ATT&CK Techniques

Possible mappings for the risks being closed:

  • T1071 — Application Layer Protocol (an exposed webhook as a reachable control surface)
  • T1078 — Valid Accounts (identity confusion in the approval keying)

🗺 Visual Investigation Diagram

Approval requested
    ↓
Old: keyed loosely, reachable webhook, unauthenticated hook
    ↓
New: keyed on sender, delivered via Telegram polling
    ↓
Policy hook authenticated + backed up
    ↓
Hook unreachable → wait, don't drop
    ↓
Poller death → alerted as an outage

⚠ Challenges

The sender-vs-message keying bug was the hardest to see, because it only shows up under concurrency — two approvals in flight at once. It took deliberately imagining an adversarial timing scenario, not just reading the code in isolation, to notice the gap.


📚 What I Learned

I learned that an approval mechanism has its own attack surface, separate from whatever it’s gating. A perfect deny-first policy behind a sloppy approval queue is still a sloppy approval queue.


➡ Next Steps

  • Re-audit the queue under deliberately concurrent test scenarios
  • Rotate and monitor the callback tokens
  • Extend the “control death is an outage” pattern to every gate in the stack
  • Review the four logged findings for whether they generalize to other queues

🧠 Reflection

Auditing the mechanism that is supposed to be the last line of defense felt like the right kind of paranoia — the gate itself needs the same scrutiny as whatever it’s protecting.


🧩 Lessons Learned

What worked

Removing a reachable webhook, authenticating the policy hook, and keying approvals on sender identity.

What broke

Message-keyed approvals under concurrency, an unauthenticated policy hook, and a poller whose death went unnoticed.

Why it broke

The queue was designed for the common case, not for concurrent or adversarial timing.

Fix / takeaway

Key authorization state on the principal, not the artifact; authenticate every hook; and alert on the death of the control itself.


📈 Skill Progression Context

This supports my cybersecurity progression because identity-scoped authorization state, attack-surface reduction, and monitoring the health of the control itself are core access-control engineering skills.


😄 TL;DR

The approval gate got its own audit — sender-keyed, webhook removed, authenticated, and alerted if it ever goes quiet.