๐Ÿ”„ Topic

Giving an agent the ability to fetch and cite real web pages means giving it a network client โ€” which means building it as if it were hostile from day one. I put together the SSRF-hardening layer for the browser-research-adapter: an allowlist, private-URL denial, robots.txt handling, content-type checks, and receipts for every fetch.


๐ŸŽฏ Goal

Let an agent fetch and extract real pages for research without letting โ€œfetch this URLโ€ become โ€œreach anything on my network I didnโ€™t intend to expose.โ€


๐Ÿ›  What I Did

I built the fetch path in layers instead of one big function.

Main areas covered:

  • a pure policy layer deciding allow/deny before any network call happens: allowlist, private-URL denial, content-type restriction, and a receipt recorded for every decision
  • a robots.txt parsing and matching layer, fixing a real bug where directives appearing before any User-agent line were leaking into every group instead of none
  • an isolated Playwright context for the actual page fetch, so the browser doing the fetching is not the browser doing anything else
  • a fetchAndCite orchestrator tying policy, robots, and fetch together into one auditable path
  • a fixture server so the guard layer could be tested against real HTTP behavior, not mocked assumptions
  • markdown extraction reusing the existing Defuddle wrapper instead of writing a second, less-tested parser

๐Ÿ”— Key Cybersecurity Connections

SSRF is what happens when โ€œfetch this URLโ€ trusts the URL. An agent that reads the web on request is a URL-fetching service with a very permissive client (me, or whatever produced the URL), so it needs the same guard a public-facing fetch endpoint would: deny private and internal address space, restrict content types, and never let policy be an afterthought bolted onto a working fetch call.

Writing policy as a separate, pure layer โ€” decide first, fetch second โ€” is what makes the guard testable and auditable independent of the browser automation underneath it.


๐Ÿ” Investigation Questions

  • Does every fetch pass through the policy layer before any connection opens?
  • What exactly counts as a private or internal address here?
  • Can robots.txt parsing be tricked into applying the wrong groupโ€™s rules?
  • Is the fetch context isolated from anything else the browser might be doing?
  • Is there a receipt for every decision, allow or deny?

๐Ÿšจ Detection Opportunities

Checks for an outbound-fetch guard:

  • a fetch reaching the network with no matching policy decision recorded
  • robots.txt directives applied to a User-agent group they were not scoped to
  • content-type restriction bypassed by a redirect or mislabeled response
  • fetch context sharing state with an unrelated browser session
  • allowlist entries added without a corresponding review record

Example:

project=browser-research-adapter
signal=network_fetch_without_policy_receipt
risk_area=ssrf_exposure
triage=trace_call_site_confirm_policy_gate_executed

๐Ÿงญ MITRE ATT&CK Techniques

Possible mappings for the risk being defended against:

  • T1090 โ€” Proxy (an unguarded fetcher can be turned into one)
  • T1595 โ€” Active Scanning (internal address space reachable through the fetcher)

๐Ÿ—บ Visual Investigation Diagram

URL requested
    โ†“
Policy layer: allowlist, private-URL check, content-type
    โ†“
Deny โ†’ receipt, stop
Allow โ†’ robots.txt check
    โ†“
Isolated fetch context
    โ†“
Extract + cite, receipt recorded

โš  Challenges

The robots.txt bug was the humbling part: directives before any User-agent line are supposed to apply to nobody, and my first pass applied them to everybody. A parsing detail that small is exactly the kind of thing that turns โ€œwe respect robots.txtโ€ into a false claim.


๐Ÿ“š What I Learned

I learned that a fetcher for an agent needs the same paranoia as a fetcher for the public internet, because from the targetโ€™s perspective there is no difference. The policy layer being pure and separate from the browser mechanics is what let me test the dangerous decisions without needing a real network.


โžก Next Steps

  • Expand the fixture serverโ€™s coverage of adversarial responses
  • Keep policy decisions and fetch mechanics in separate, independently testable layers
  • Review the allowlist on a schedule, not just at creation
  • Watch for this pattern reused anywhere else an agent gets a network client

๐Ÿง  Reflection

Building this felt like writing the control I would want to find already in place during an audit โ€” deny by default, decide before you connect, and leave a receipt.


๐Ÿงฉ Lessons Learned

What worked

Separating the pure policy decision from the browser mechanics that execute it.

What broke

Robots.txt directives before any User-agent line applied globally instead of to nobody.

Why it broke

The parser assumed a group context existed before the file declared one.

Fix / takeaway

Test parsing against the specificationโ€™s edge cases, not just the common case โ€” and keep every network-facing decision behind a pure, auditable gate.


๐Ÿ“ˆ Skill Progression Context

This supports my cybersecurity progression because SSRF defense, input validation, and policy-before-action design are core web application security skills, practiced here on a client I built and can fully inspect.


๐Ÿ˜„ TL;DR

Gave the agent a web fetcher, and made it prove every request is safe before it ever leaves the house.