🧪 Lab 11 – Assessing a Candidate Tool Before Adoption
Lab Objective
Practice a security-first evaluation of a third-party command-line tool before adopting it:
- pin the source being examined
- inspect declared capabilities and installation behavior
- compare it with an existing manual workflow
- run a bounded trial with no credentials or background service
- document a keep, defer, or reject decision and rollback path
This lab uses a disposable checkout. Do not run an unfamiliar installer with elevated privileges.
Lab Environment
- System: Linux or macOS with Git and a shell
- Candidate: a public command-line repository chosen for review
- Working directory: a disposable directory under
/tmp - Permissions: normal user only; no tokens, production data, or browser profiles
- Boundary: no daemon, scheduled job, shell hook, or automatic integration
Scenario
A tool promises to speed up a local research task. Before it touches any active workflow, determine what it can access, whether it duplicates existing capability, and whether the benefit justifies the new trust boundary.
Commands Practiced
| Command | Purpose |
|---|---|
git ls-remote |
Identify the remote revision being reviewed |
git clone --depth 1 |
Make a small disposable checkout |
git rev-parse HEAD |
Record the exact checked-out revision |
rg |
Search the source for install hooks and persistent behavior |
find |
Inspect the repository layout and manifests |
git diff --check |
Check local documentation for whitespace errors |
Step 1 - Create a Disposable Review Area
mkdir -p /tmp/tool-review && cd /tmp/tool-review
git ls-remote <repository-url> HEAD
git clone --depth 1 <repository-url> candidate-tool
git -C candidate-tool rev-parse HEAD
Record the final commit identifier in your notes. A repository name alone is not enough: its contents can change over time.
Step 2 - Inspect the Capability Surface
Search for common signs that a tool installs persistent behavior or expects sensitive access:
rg -n -i "postinstall|preinstall|launchagent|launchd|daemon|watch|oauth|token|api.?key|mcp" candidate-tool
find candidate-tool -maxdepth 2 -type f | sort | sed -n '1,120p'
Read the manifest, installation instructions, and configuration examples. A match is not proof of a problem; it is a reason to understand the code path before proceeding.
Step 3 - Compare with a Manual Baseline
Write down the exact task the candidate claims to improve. Perform the task once with existing tools, then state what measurable improvement would justify the addition.
Example decision table:
| Question | Baseline | Candidate result | Decision |
|---|---|---|---|
| Can the result be verified from source? | Yes | Yes | Continue review |
| Does it need a credential? | No | No | Continue review |
| Does it leave a background process? | No | No | Continue review |
| Does it reduce real verification effort? | N/A | Not demonstrated | Defer |
Fast output without a verification benefit is not enough for adoption.
Step 4 - Run Only a Bounded Trial
Use a test file or a non-sensitive fixture. Do not point the candidate at an entire home directory, browser profile, or production repository.
Before and after the trial, check for unexpected state:
ps aux | rg -i "candidate-tool|node|python"
find "$HOME" -maxdepth 2 -type d -name '*candidate*' 2>/dev/null
If the tool needs a service, credentials, or broad filesystem access before it can demonstrate value, stop the trial and record that as a finding.
Step 5 - Make the Decision and Record Rollback
Choose one outcome:
- Keep manually: a narrow, useful command with documented bounds.
- Defer: promising, but no measured advantage yet.
- Reject: duplicated capability, unclear provenance, excessive access, or unacceptable persistence.
For a kept tool, record the revision, permissions, input/output boundary, verification method, and removal command. For a rejected tool, remove the disposable checkout and any trial cache.
Security Takeaways
- Installation is a trust decision. Treat it as a change with a source, scope, and rollback path.
- Pin what you review. A repository URL is not a stable artifact.
- Bounded trials reveal behavior. Test the smallest useful task before connecting a tool to real data.
- Capabilities have operating costs. Background services, credentials, and retained state expand the attack surface.
- Rejection can be the right outcome. A documented no is evidence that the boundary was considered.
