πŸ”„ Topic

Testing a fully autonomous local agent (Hermes) on a real app build, plus wiring it to a Telegram control channel.


🎯 Goal

Measure what unsupervised agent autonomy is actually worth: where it shines, where it fails, and what controls the failures demand.


πŸ›  What I Did

I benchmarked the Hermes agent by giving it complete, written briefs for a travel packing app β€” redesign, gear kits with quantities, weight estimates, PWA install support β€” and letting it execute end to end in batches with a git commit after each verified batch. I also installed the agent gateway as a launchd service and connected a Telegram channel for remote tasking. The honest verdict from the benchmark: strong one-shot building, unreliable self-debugging. When its own code broke, the agent often could not diagnose itself, and a stronger model or a human had to step in.

Main areas covered:

  • written briefs as executable task specifications
  • batch execution with commit checkpoints
  • PWA build: manifest, service worker, install support
  • launchd service for the agent gateway
  • Telegram as a remote control channel
  • documenting benchmark results honestly

πŸ”— Key Cybersecurity Connections

Autonomous agents that commit code and run as system services are privileged automation. Commit checkpoints are rollback points, the launchd service is persistence by design, and a Telegram control channel is remote access that must be scoped and authenticated like any other.


πŸ” Investigation Questions

  • What is the agent’s actual failure mode when code breaks?
  • Does every batch have a verifiable checkpoint?
  • Who can send commands through the Telegram channel?
  • What does the launchd service run with, and as whom?
  • Are the agent’s commits reviewable and attributable?

🚨 Detection Opportunities

Potential monitoring ideas:

  • commits by the agent outside a briefed task
  • launchd service configuration changes
  • commands from unknown Telegram accounts
  • repeated failed self-debug loops
  • scope creep beyond the written brief

Example:

project=hermes-agent-benchmark
change_type=autonomous_commit_batch
risk_area=unattended_code_execution
triage=review_commits_against_brief_scope

🧭 MITRE ATT&CK Techniques

Possible mappings depending on confirmed behavior:

  • T1543 β€” Create or Modify System Process
  • T1102 β€” Web Service
  • T1059 β€” Command and Scripting Interpreter

πŸ—Ί Visual Investigation Diagram

Written brief
    ↓ Autonomous batches
    ↓ Commit checkpoints
    ↓ Benchmark verdict
    ↓ Controls for the weak spots

⚠ Challenges

The challenge was accepting the result. I wanted the agent to be fully autonomous; the evidence said it builds well but cannot reliably fix itself. Writing that down honestly mattered more than the wish.


πŸ“š What I Learned

I learned that autonomy has a measurable shape: this agent is a strong builder and a weak debugger. That asymmetry dictates the controls β€” checkpoints, verification, and an escalation path β€” rather than blind trust.


➑ Next Steps

  • Keep commit-per-batch as a hard rule
  • Restrict and authenticate the Telegram channel
  • Route self-debug failures to a stronger model automatically
  • Re-run the benchmark after each agent upgrade

🧠 Reflection

This was useful because benchmarking my own automation felt like red-teaming my own trust: I now know exactly where the agent breaks, which is where a defender’s attention belongs.


🧩 Lessons Learned

What worked

Complete written briefs and commit checkpoints after every batch.

What broke

The agent’s self-debugging when its own build failed.

Why it broke

Diagnosing failures needs stronger reasoning than one-shot generation.

Fix / takeaway

Design for the failure mode: checkpoints, verification, and escalation instead of assumed autonomy.


πŸ“ˆ Skill Progression Context

This supports my cybersecurity progression because evaluating automation honestly β€” capabilities, failure modes, persistence, remote access β€” is exactly how defenders must assess the agentic systems arriving in every organization.


πŸ˜„ TL;DR

Great builder, bad self-mechanic; plan controls accordingly.