๐ Day 182 โ Building an SSRF-Safe Fetcher for an Agent That Reads the Web
๐ Topic
Giving an agent the ability to fetch and cite real web pages means giving it a network client โ which means building it as if it were hostile from day one. I put together the SSRF-hardening layer for the browser-research-adapter: an allowlist, private-URL denial, robots.txt handling, content-type checks, and receipts for every fetch.
๐ฏ Goal
Let an agent fetch and extract real pages for research without letting โfetch this URLโ become โreach anything on my network I didnโt intend to expose.โ
๐ What I Did
I built the fetch path in layers instead of one big function.
Main areas covered:
- a pure policy layer deciding allow/deny before any network call happens: allowlist, private-URL denial, content-type restriction, and a receipt recorded for every decision
- a robots.txt parsing and matching layer, fixing a real bug where directives appearing before any User-agent line were leaking into every group instead of none
- an isolated Playwright context for the actual page fetch, so the browser doing the fetching is not the browser doing anything else
- a
fetchAndCiteorchestrator tying policy, robots, and fetch together into one auditable path - a fixture server so the guard layer could be tested against real HTTP behavior, not mocked assumptions
- markdown extraction reusing the existing Defuddle wrapper instead of writing a second, less-tested parser
๐ Key Cybersecurity Connections
SSRF is what happens when โfetch this URLโ trusts the URL. An agent that reads the web on request is a URL-fetching service with a very permissive client (me, or whatever produced the URL), so it needs the same guard a public-facing fetch endpoint would: deny private and internal address space, restrict content types, and never let policy be an afterthought bolted onto a working fetch call.
Writing policy as a separate, pure layer โ decide first, fetch second โ is what makes the guard testable and auditable independent of the browser automation underneath it.
๐ Investigation Questions
- Does every fetch pass through the policy layer before any connection opens?
- What exactly counts as a private or internal address here?
- Can robots.txt parsing be tricked into applying the wrong groupโs rules?
- Is the fetch context isolated from anything else the browser might be doing?
- Is there a receipt for every decision, allow or deny?
๐จ Detection Opportunities
Checks for an outbound-fetch guard:
- a fetch reaching the network with no matching policy decision recorded
- robots.txt directives applied to a User-agent group they were not scoped to
- content-type restriction bypassed by a redirect or mislabeled response
- fetch context sharing state with an unrelated browser session
- allowlist entries added without a corresponding review record
Example:
project=browser-research-adapter
signal=network_fetch_without_policy_receipt
risk_area=ssrf_exposure
triage=trace_call_site_confirm_policy_gate_executed
๐งญ MITRE ATT&CK Techniques
Possible mappings for the risk being defended against:
- T1090 โ Proxy (an unguarded fetcher can be turned into one)
- T1595 โ Active Scanning (internal address space reachable through the fetcher)
๐บ Visual Investigation Diagram
URL requested
โ
Policy layer: allowlist, private-URL check, content-type
โ
Deny โ receipt, stop
Allow โ robots.txt check
โ
Isolated fetch context
โ
Extract + cite, receipt recorded
โ Challenges
The robots.txt bug was the humbling part: directives before any User-agent line are supposed to apply to nobody, and my first pass applied them to everybody. A parsing detail that small is exactly the kind of thing that turns โwe respect robots.txtโ into a false claim.
๐ What I Learned
I learned that a fetcher for an agent needs the same paranoia as a fetcher for the public internet, because from the targetโs perspective there is no difference. The policy layer being pure and separate from the browser mechanics is what let me test the dangerous decisions without needing a real network.
โก Next Steps
- Expand the fixture serverโs coverage of adversarial responses
- Keep policy decisions and fetch mechanics in separate, independently testable layers
- Review the allowlist on a schedule, not just at creation
- Watch for this pattern reused anywhere else an agent gets a network client
๐ง Reflection
Building this felt like writing the control I would want to find already in place during an audit โ deny by default, decide before you connect, and leave a receipt.
๐งฉ Lessons Learned
What worked
Separating the pure policy decision from the browser mechanics that execute it.
What broke
Robots.txt directives before any User-agent line applied globally instead of to nobody.
Why it broke
The parser assumed a group context existed before the file declared one.
Fix / takeaway
Test parsing against the specificationโs edge cases, not just the common case โ and keep every network-facing decision behind a pure, auditable gate.
๐ Skill Progression Context
This supports my cybersecurity progression because SSRF defense, input validation, and policy-before-action design are core web application security skills, practiced here on a client I built and can fully inspect.
๐ TL;DR
Gave the agent a web fetcher, and made it prove every request is safe before it ever leaves the house.
