🧪 Lab 14 – Building and Breaking an SSRF-Safe Fetch Guard
Lab Objective
Build a minimal outbound-fetch guard, then deliberately attack it, using the same reasoning that found two live bypasses in a real IP-pinning fix:
- implement a policy layer that denies private/internal address space before any request
- implement DNS-safe connection logic that re-validates the address at connection time, not just at the URL-parsing stage
- reproduce a time-of-check/time-of-use (DNS rebinding) bypass against a naive version of the guard
- reproduce an IP-literal bypass against a guard that only checks hostnames
- fix both, and prove the fix with a live local target
This mirrors real findings from hardening an agent’s web-fetching tool: a guard that checks the URL string is not the same as a guard that checks what the request actually connects to.
Lab Environment
- Language: Python 3.10+ (standard library
socket,http.server,ipaddress) - Target: a local HTTP server bound to
127.0.0.1, standing in for an internal resource that should never be reachable through the fetcher - No third-party dependencies required — everything here uses the standard library so it runs anywhere
All commands assume a scratch directory; nothing here touches a real network beyond localhost.
Scenario
An agent has a tool: “fetch this URL and summarize it.” The URL comes from an untrusted source — a web page, a user request, anywhere. The fetcher must never be tricked into reaching 127.0.0.1, 169.254.169.254 (the cloud metadata range), or any other private address, no matter how the request is phrased.
Commands / Code Practiced
| Tool | Purpose |
|---|---|
ipaddress.ip_address(...).is_private |
Check whether a resolved address is private/internal |
socket.getaddrinfo(host, port) |
Resolve a hostname to its actual connecting address |
http.server.HTTPServer |
Stand up a local target to attack |
custom is_safe_url(url) |
The policy layer under test |
Step 1 - Stand Up the “Internal” Target
# server.py
from http.server import HTTPServer, BaseHTTPRequestHandler
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
self.send_response(200)
self.end_headers()
self.wfile.write(b"INTERNAL SECRET: should never be fetched")
HTTPServer(("127.0.0.1", 8765), Handler).serve_forever()
Run it in one terminal: python3 server.py
Step 2 - A Naive Guard (Hostname-Only Check)
# naive_guard.py
import ipaddress
from urllib.parse import urlparse
def is_safe_url_naive(url: str) -> bool:
host = urlparse(url).hostname
try:
ipaddress.ip_address(host)
return False # looks like a literal IP -> deny... but only if it parses
except ValueError:
return True # not a literal IP string -> "safe" (WRONG)
This is deliberately broken the way Defect 3 was broken: it only catches an IP address written as a literal string, and only rejects it — it never actually validates what a hostname resolves to.
python3 -c "
from naive_guard import is_safe_url_naive
print(is_safe_url_naive('http://localhost:8765/')) # True -- WRONG, this is unsafe
print(is_safe_url_naive('http://127.0.0.1:8765/')) # False -- caught, but only by luck of string matching
"
localhost sails through. Fetch it with any plain HTTP client and the “internal secret” comes back — the naive guard never actually resolved and checked the address.
Step 3 - Fix It: Resolve, Then Validate
# safe_guard.py
import socket
import ipaddress
from urllib.parse import urlparse
def is_safe_url(url: str) -> bool:
parsed = urlparse(url)
host = parsed.hostname
port = parsed.port or (443 if parsed.scheme == "https" else 80)
if not host:
return False
try:
infos = socket.getaddrinfo(host, port)
except socket.gaierror:
return False
for family, _, _, _, sockaddr in infos:
addr = sockaddr[0]
ip = ipaddress.ip_address(addr)
if ip.is_private or ip.is_loopback or ip.is_link_local or ip.is_reserved:
return False # deny on ANY resolved address being unsafe
return True
python3 -c "
from safe_guard import is_safe_url
print(is_safe_url('http://localhost:8765/')) # False -- correctly denied
print(is_safe_url('http://127.0.0.1:8765/')) # False -- correctly denied
print(is_safe_url('http://example.com/')) # True
"
Now localhost resolves to 127.0.0.1, which is checked and denied — the guard validates the real address, not the string.
Step 4 - The Remaining Gap: Check-Then-Connect Is Still Not Check-At-Connect
Even this fixed guard has the real-world DNS-rebinding gap: is_safe_url resolves the host once to check it, and the actual HTTP client resolves it again, moments later, to connect. For a normal DNS name that never changes, this is fine. For an attacker-controlled DNS name with a very short TTL, the second resolution can legitimately return a different address than the first.
# simulate why re-resolution matters: two calls, two different results possible
import socket
print(socket.getaddrinfo("example.com", 80)[0][4][0]) # check-time address
# ... time passes, attacker's DNS TTL expires ...
print(socket.getaddrinfo("example.com", 80)[0][4][0]) # connect-time address -- could differ for a hostile domain
The real fix (as in the production incident) is to pin the address you validated and force the connection to use that exact address — never re-resolve between check and connect. That is what Chromium’s --host-resolver-rules pinning does, and what the lab’s safe_guard.py does not yet do — left as the natural next exercise.
Step 5 - Prove the Fix Blocks the Real Target
curl --resolve localhost:8765:127.0.0.1 http://localhost:8765/ # what an unguarded client would happily return
python3 -c "
from safe_guard import is_safe_url
assert not is_safe_url('http://localhost:8765/')
print('BLOCKED as expected')
"
Security Takeaways
- Check what you connect to, not what the URL says. String-based checks (
naive_guard.py) catch typing conventions, not addresses. - Resolve before you validate.
localhost, custom DNS entries, and short hostnames all deserve real resolution before a decision is made. - Check-then-connect has a gap of its own. Any guard that resolves once to validate and again to connect is exposed to DNS rebinding — pin the validated address if the target matters.
- Deny broadly. Private, loopback, link-local, and reserved ranges all need denial — a guard that only blocks
127.0.0.1misses169.254.169.254. - Prove it against a live target. A guard that “looks right” in review can still be wrong; run it against a real local server before trusting it.
Where This Applies Beyond the Lab
This is the exact reasoning behind real SSRF defenses in any service that fetches user-supplied URLs: webhook receivers, link-preview generators, PDF renderers, and — as this lab is modeled on — an AI agent’s web-research tool. The pattern (resolve, validate the real address, pin it, deny broad ranges) transfers directly.
