Lab Objective

Practice extracting readable text from a local HTML file while keeping the operation narrow and inspectable:

  • use one explicit local input
  • enforce a timeout and output cap in the wrapper
  • separate extracted content from source metadata
  • verify the wrapper before use
  • avoid automatic ingestion, browser profiles, and persistent services

Lab Environment

  • System: Linux or macOS shell
  • Input: a non-sensitive local HTML file
  • Tool: a local extraction wrapper, for example scripts/defuddle-extract
  • Output: Markdown on standard output; source metadata on standard error
  • Boundary: one-shot command only; no automatic index or knowledge-base import

Scenario

You need the main text from a saved HTML page for a note. A full browser automation stack would have access to much more state than this task needs, so the extraction must remain a small, visible command with an explicit input and an inspectable output.


Commands Practiced

Command Purpose
zsh -n Check shell syntax without running the wrapper
python3 Run focused wrapper tests when supplied by the project
wc -c Inspect output size
shasum -a 256 Record an integrity hash for the extracted result
sed -n Review a small, readable sample of the output

Step 1 - Verify the Wrapper Before Running It

Inspect the wrapper’s stated limits and check its syntax:

zsh -n scripts/defuddle-extract
python3 scripts/test_defuddle_extract.py

The wrapper should make its boundaries visible: acceptable input type, timeout, output cap, and where it writes data.


Step 2 - Extract One Local File

Use a disposable fixture or a saved page that contains no sensitive information:

scripts/defuddle-extract ./fixtures/example.html \
  > extracted.md \
  2> source-metadata.log

This separates the readable result from details about the source. Keeping those streams separate makes later review easier.


Step 3 - Inspect Size, Content, and Integrity

wc -c extracted.md source-metadata.log
sed -n '1,80p' extracted.md
cat source-metadata.log
shasum -a 256 extracted.md

Check that the result is within the expected size limit, contains the intended text, and names the source accurately. The hash is useful when a note needs to record exactly which extracted artifact was reviewed.


Step 4 - Check for Unwanted State

The command should exit when extraction ends. Confirm that it did not create a browser profile, a long-lived service, or an automatic import path:

ps aux | rg -i "defuddle|chrome|playwright" 
find . -maxdepth 2 -type f | sort

Expected result: the extracted Markdown and metadata files are visible, and no unrelated background process remains.


Step 5 - Decide What Happens Next

Review the Markdown before it is quoted, summarized, or placed in a note. The extraction command ends here; any later storage, indexing, or sharing step should be a separate decision.


Security Takeaways

  1. Explicit input limits collection. One named file is easier to reason about than a browser session.
  2. Timeouts and caps are security controls. They bound resource use and accidental bulk output.
  3. Metadata preserves provenance. Store where text came from separately from the text itself.
  4. One-shot tooling is easier to audit. Short-lived processes leave a smaller state footprint.
  5. Extraction is not authorization to ingest. Review the result before any later workflow consumes it.