🧪 Lab 12 – Extracting Text from a Local HTML File Safely
Lab Objective
Practice extracting readable text from a local HTML file while keeping the operation narrow and inspectable:
- use one explicit local input
- enforce a timeout and output cap in the wrapper
- separate extracted content from source metadata
- verify the wrapper before use
- avoid automatic ingestion, browser profiles, and persistent services
Lab Environment
- System: Linux or macOS shell
- Input: a non-sensitive local HTML file
- Tool: a local extraction wrapper, for example
scripts/defuddle-extract - Output: Markdown on standard output; source metadata on standard error
- Boundary: one-shot command only; no automatic index or knowledge-base import
Scenario
You need the main text from a saved HTML page for a note. A full browser automation stack would have access to much more state than this task needs, so the extraction must remain a small, visible command with an explicit input and an inspectable output.
Commands Practiced
| Command | Purpose |
|---|---|
zsh -n |
Check shell syntax without running the wrapper |
python3 |
Run focused wrapper tests when supplied by the project |
wc -c |
Inspect output size |
shasum -a 256 |
Record an integrity hash for the extracted result |
sed -n |
Review a small, readable sample of the output |
Step 1 - Verify the Wrapper Before Running It
Inspect the wrapper’s stated limits and check its syntax:
zsh -n scripts/defuddle-extract
python3 scripts/test_defuddle_extract.py
The wrapper should make its boundaries visible: acceptable input type, timeout, output cap, and where it writes data.
Step 2 - Extract One Local File
Use a disposable fixture or a saved page that contains no sensitive information:
scripts/defuddle-extract ./fixtures/example.html \
> extracted.md \
2> source-metadata.log
This separates the readable result from details about the source. Keeping those streams separate makes later review easier.
Step 3 - Inspect Size, Content, and Integrity
wc -c extracted.md source-metadata.log
sed -n '1,80p' extracted.md
cat source-metadata.log
shasum -a 256 extracted.md
Check that the result is within the expected size limit, contains the intended text, and names the source accurately. The hash is useful when a note needs to record exactly which extracted artifact was reviewed.
Step 4 - Check for Unwanted State
The command should exit when extraction ends. Confirm that it did not create a browser profile, a long-lived service, or an automatic import path:
ps aux | rg -i "defuddle|chrome|playwright"
find . -maxdepth 2 -type f | sort
Expected result: the extracted Markdown and metadata files are visible, and no unrelated background process remains.
Step 5 - Decide What Happens Next
Review the Markdown before it is quoted, summarized, or placed in a note. The extraction command ends here; any later storage, indexing, or sharing step should be a separate decision.
Security Takeaways
- Explicit input limits collection. One named file is easier to reason about than a browser session.
- Timeouts and caps are security controls. They bound resource use and accidental bulk output.
- Metadata preserves provenance. Store where text came from separately from the text itself.
- One-shot tooling is easier to audit. Short-lived processes leave a smaller state footprint.
- Extraction is not authorization to ingest. Review the result before any later workflow consumes it.
