📅 Day 148 — Building Least-Privilege Tool Profiles for My AI Agents (and Finding My Own Risk Scorer Was Inverted)
🔄 Topic
My AI agents can currently reach every tool I’ve connected — filesystem, Gmail, GitHub, everything — regardless of what the task actually needs. I started building scoped permission profiles so each task type only gets the access it requires, then measured the risk of each profile and caught a real bug in my own scoring logic.
🎯 Goal
Replace “one agent, every tool, all the time” with named profiles that grant only what a given task type actually needs — and be able to justify, in numbers, why each profile is safer or riskier than another.
🛠 What I Did
I designed six task-scoped profiles, documented an explicit allowlist and safety rules for each, found and fixed a real over-broad tool bundle, and then built a scoring pass to measure each profile’s actual risk instead of assuming it from the design docs.
Main areas covered:
- six profiles: coding-core, coding-github, autonomous-builder, research-handoff, soc-lab, personal-admin — each with its own allowed-tools list and safety rules
- auditing the existing tool infrastructure and finding that one bundled profile granted both filesystem and Gmail access together, when two of my six new profiles specifically needed one or the other but never both
- splitting that bundle into two separate, additive profiles (filesystem-only, Gmail-only) instead of leaving two profiles with a documented-but-unfixed gap
- building a measurement pass that scores each profile on tool surface, permission level, personal-data exposure, and unattended-run risk
- catching a real bug where the personal-data-risk scorer matched “gmail”/”calendar” inside a profile’s own forbidden list, which meant every profile that explicitly denied Gmail access was scored as if it had Gmail access
🔗 Key Cybersecurity Connections
This is least-privilege access control applied to AI agents instead of human accounts: an autonomous coding agent doesn’t need email access, and a personal-admin assistant doesn’t need repository access, so neither should default to having both just because one convenient bundle already existed. The scoring bug is the more interesting lesson — a risk-scoring rule that checks for the presence of a keyword instead of the presence of a granted capability will happily score “explicitly denies X” the same as “has X.” A control that inverts its own signal is worse than no control, because it looks like coverage while measuring the opposite of what it claims to.
🔍 Investigation Questions
- Does each profile’s actual tool grant match what the task genuinely needs, or is it inherited from a convenient existing bundle?
- Is any profile one bundled tool config away from silently getting access it shouldn’t have?
- Does the risk scorer measure granted capability, or does it pattern-match on text that could appear in either a grant or a denial?
- Which profiles are safe enough to run daily, and which should be rare, deliberate, and reviewed each time?
- Is a “gap” documented and left open, or documented and actually closed?
🚨 Detection Opportunities
Potential monitoring ideas:
- a profile whose granted tools silently expand after a shared bundle changes
- a risk score that doesn’t change even after tool grants change (sign the scorer isn’t reading real state)
- any profile combining filesystem write access with unattended/autonomous execution
- new profiles created without a corresponding risk-scoring pass
Example:
project=agent-mcp-profiles
change_type=least_privilege_profile_rollout
risk_area=tool_overprivilege
triage=diff_granted_tools_against_documented_need
🧭 MITRE ATT&CK Techniques
Framed as the defensive control this work targets:
- T1078.004 — Valid Accounts: Cloud Accounts (the analog here is an over-privileged agent profile acting like an over-privileged account)
- T1526 — Cloud Service Discovery (the same reasoning applies to auditing what an agent profile can reach before assuming it’s scoped correctly)
🗺 Visual Investigation Diagram
One agent, every tool, always
↓
Design six task-scoped profiles
↓
Audit existing tool bundles against real profile needs
↓
Split over-broad bundle (fs+gmail) into fs-only / gmail-only
↓
Build risk scorer
↓
Scorer bug found: forbidden-list keywords inflated risk score
↓
Fix scorer, re-verify only the actually-risky profile scores high
⚠ Challenges
The scoring bug was the kind that looks like it’s working — every profile produced a number, the numbers seemed plausible, and nothing crashed. It only became obvious as wrong because every single coding profile scored high on personal-data risk, which didn’t match what those profiles actually granted. A quiet, plausible-looking wrong number is more dangerous than an obvious failure.
📚 What I Learned
I learned that a risk score is only as trustworthy as what it actually reads. Pattern-matching on a keyword instead of resolving actual granted capability produces a number that looks like a measurement but is actually measuring the wrong thing — and it will keep producing confident, wrong numbers indefinitely unless something forces a sanity check against what the profiles really do.
➡ Next Steps
- Move from documented profiles to real, enforced tool-list restriction on actual agent runs
- Re-run the risk scoring pass after any profile or Docker MCP bundle change
- Decide the fate of the local-model backend currently blocking one deferred reconciliation, then revisit that gap
🧠 Reflection
Building the profiles felt like the real work; finding the scoring bug turned out to matter more, because a wrong measurement of risk is worse than no measurement — it creates false confidence exactly where scrutiny should be highest.
🧩 Lessons Learned
What worked
Splitting the over-broad filesystem+Gmail bundle into two additive, reversible profiles instead of leaving a documented gap unfixed.
What broke
The personal-data-risk scorer, which matched keywords inside a profile’s forbidden list and scored denial the same as access.
Why it broke
Text pattern-matching isn’t the same as resolving actual granted capability — a rule that checks for a word will fire whether that word appears in a grant or a refusal.
Fix / takeaway
Any automated risk or compliance scorer needs to be checked against known-good and known-bad cases before its numbers are trusted, the same as any other piece of security tooling.
📈 Skill Progression Context
This supports my cybersecurity progression because designing least-privilege profiles and then catching a false-confidence bug in my own risk-scoring logic is a direct, hands-on rehearsal of the access-control and control-validation thinking real security tooling review requires.
😄 TL;DR
Built six least-privilege tool profiles for my AI agents, split an over-broad bundle in two, and caught my own risk scorer quietly grading “explicitly forbidden” the same as “granted.”
