Lab Objective

Build a minimal outbound-fetch guard, then deliberately attack it, using the same reasoning that found two live bypasses in a real IP-pinning fix:

  • implement a policy layer that denies private/internal address space before any request
  • implement DNS-safe connection logic that re-validates the address at connection time, not just at the URL-parsing stage
  • reproduce a time-of-check/time-of-use (DNS rebinding) bypass against a naive version of the guard
  • reproduce an IP-literal bypass against a guard that only checks hostnames
  • fix both, and prove the fix with a live local target

This mirrors real findings from hardening an agent’s web-fetching tool: a guard that checks the URL string is not the same as a guard that checks what the request actually connects to.


Lab Environment

  • Language: Python 3.10+ (standard library socket, http.server, ipaddress)
  • Target: a local HTTP server bound to 127.0.0.1, standing in for an internal resource that should never be reachable through the fetcher
  • No third-party dependencies required — everything here uses the standard library so it runs anywhere

All commands assume a scratch directory; nothing here touches a real network beyond localhost.


Scenario

An agent has a tool: “fetch this URL and summarize it.” The URL comes from an untrusted source — a web page, a user request, anywhere. The fetcher must never be tricked into reaching 127.0.0.1, 169.254.169.254 (the cloud metadata range), or any other private address, no matter how the request is phrased.


Commands / Code Practiced

Tool Purpose
ipaddress.ip_address(...).is_private Check whether a resolved address is private/internal
socket.getaddrinfo(host, port) Resolve a hostname to its actual connecting address
http.server.HTTPServer Stand up a local target to attack
custom is_safe_url(url) The policy layer under test

Step 1 - Stand Up the “Internal” Target

# server.py
from http.server import HTTPServer, BaseHTTPRequestHandler

class Handler(BaseHTTPRequestHandler):
    def do_GET(self):
        self.send_response(200)
        self.end_headers()
        self.wfile.write(b"INTERNAL SECRET: should never be fetched")

HTTPServer(("127.0.0.1", 8765), Handler).serve_forever()

Run it in one terminal: python3 server.py


Step 2 - A Naive Guard (Hostname-Only Check)

# naive_guard.py
import ipaddress
from urllib.parse import urlparse

def is_safe_url_naive(url: str) -> bool:
    host = urlparse(url).hostname
    try:
        ipaddress.ip_address(host)
        return False   # looks like a literal IP -> deny... but only if it parses
    except ValueError:
        return True     # not a literal IP string -> "safe" (WRONG)

This is deliberately broken the way Defect 3 was broken: it only catches an IP address written as a literal string, and only rejects it — it never actually validates what a hostname resolves to.

python3 -c "
from naive_guard import is_safe_url_naive
print(is_safe_url_naive('http://localhost:8765/'))   # True -- WRONG, this is unsafe
print(is_safe_url_naive('http://127.0.0.1:8765/'))    # False -- caught, but only by luck of string matching
"

localhost sails through. Fetch it with any plain HTTP client and the “internal secret” comes back — the naive guard never actually resolved and checked the address.


Step 3 - Fix It: Resolve, Then Validate

# safe_guard.py
import socket
import ipaddress
from urllib.parse import urlparse

def is_safe_url(url: str) -> bool:
    parsed = urlparse(url)
    host = parsed.hostname
    port = parsed.port or (443 if parsed.scheme == "https" else 80)
    if not host:
        return False
    try:
        infos = socket.getaddrinfo(host, port)
    except socket.gaierror:
        return False
    for family, _, _, _, sockaddr in infos:
        addr = sockaddr[0]
        ip = ipaddress.ip_address(addr)
        if ip.is_private or ip.is_loopback or ip.is_link_local or ip.is_reserved:
            return False   # deny on ANY resolved address being unsafe
    return True

python3 -c "
from safe_guard import is_safe_url
print(is_safe_url('http://localhost:8765/'))    # False -- correctly denied
print(is_safe_url('http://127.0.0.1:8765/'))     # False -- correctly denied
print(is_safe_url('http://example.com/'))        # True
"

Now localhost resolves to 127.0.0.1, which is checked and denied — the guard validates the real address, not the string.


Step 4 - The Remaining Gap: Check-Then-Connect Is Still Not Check-At-Connect

Even this fixed guard has the real-world DNS-rebinding gap: is_safe_url resolves the host once to check it, and the actual HTTP client resolves it again, moments later, to connect. For a normal DNS name that never changes, this is fine. For an attacker-controlled DNS name with a very short TTL, the second resolution can legitimately return a different address than the first.

# simulate why re-resolution matters: two calls, two different results possible
import socket
print(socket.getaddrinfo("example.com", 80)[0][4][0])  # check-time address
# ... time passes, attacker's DNS TTL expires ...
print(socket.getaddrinfo("example.com", 80)[0][4][0])  # connect-time address -- could differ for a hostile domain

The real fix (as in the production incident) is to pin the address you validated and force the connection to use that exact address — never re-resolve between check and connect. That is what Chromium’s --host-resolver-rules pinning does, and what the lab’s safe_guard.py does not yet do — left as the natural next exercise.


Step 5 - Prove the Fix Blocks the Real Target

curl --resolve localhost:8765:127.0.0.1 http://localhost:8765/   # what an unguarded client would happily return

python3 -c "
from safe_guard import is_safe_url
assert not is_safe_url('http://localhost:8765/')
print('BLOCKED as expected')
"

Security Takeaways

  1. Check what you connect to, not what the URL says. String-based checks (naive_guard.py) catch typing conventions, not addresses.
  2. Resolve before you validate. localhost, custom DNS entries, and short hostnames all deserve real resolution before a decision is made.
  3. Check-then-connect has a gap of its own. Any guard that resolves once to validate and again to connect is exposed to DNS rebinding — pin the validated address if the target matters.
  4. Deny broadly. Private, loopback, link-local, and reserved ranges all need denial — a guard that only blocks 127.0.0.1 misses 169.254.169.254.
  5. Prove it against a live target. A guard that “looks right” in review can still be wrong; run it against a real local server before trusting it.

Where This Applies Beyond the Lab

This is the exact reasoning behind real SSRF defenses in any service that fetches user-supplied URLs: webhook receivers, link-preview generators, PDF renderers, and — as this lab is modeled on — an AI agent’s web-research tool. The pattern (resolve, validate the real address, pin it, deny broad ranges) transfers directly.