πŸ”„ Topic

Building and benchmarking the local LLM coding stack β€” a test harness, routing rules, and fixing what code review found.


🎯 Goal

Know, with evidence, which local model to trust for which coding task β€” and clean up the reliability bugs review surfaced.


πŸ›  What I Did

I documented and benchmarked the local coding stack: a harness for running models against real tasks, routing rules for which model handles what, and workflow docs so the setup survives model swaps. Then I ran a review pass over my own orchestration code and fixed what it found: a routing fallback that could pick the wrong provider, unsafe provider defaults, a cancellation path that could race, and orphaned processes left behind after failures. I also bootstrapped a proper operating layer for the workstation’s coding agents.

Main areas covered:

  • benchmark harness for local coding models
  • routing rules based on measured capability
  • workflow documentation for model swaps
  • routing fallback and provider default fixes
  • cancellation safety and orphan process cleanup
  • operating layer bootstrap for coding agents

πŸ”— Key Cybersecurity Connections

Reliability bugs are security bugs waiting for context: an unsafe fallback is an unintended trust decision, a cancellation race is a state-integrity flaw, and orphaned processes are unaccounted-for execution. Reviewing my own automation like hostile code is defensive practice.


πŸ” Investigation Questions

  • Which model does each task type actually route to, and why?
  • What happens when the preferred provider is unavailable?
  • Can a cancelled task leave work half-applied?
  • Are there processes running that no active task owns?
  • Do the benchmark results justify the routing table?

🚨 Detection Opportunities

Potential monitoring ideas:

  • fallback routing activating unexpectedly
  • orphaned agent processes accumulating
  • tasks cancelled but still producing output
  • provider configuration drift from defaults
  • benchmark regressions after model updates

Example:

project=local-coding-stack
change_type=orchestration_reliability_fix
risk_area=unintended_execution_paths
triage=audit_fallbacks_cancellation_and_orphans

🧭 MITRE ATT&CK Techniques

Possible mappings depending on confirmed behavior:

  • T1059 β€” Command and Scripting Interpreter
  • T1057 β€” Process Discovery
  • T1565 β€” Data Manipulation

πŸ—Ί Visual Investigation Diagram

Benchmark harness
    ↓ Measured capability
    ↓ Routing rules
    ↓ Review findings
    ↓ Fixed, documented stack

⚠ Challenges

The challenge was reviewing my own orchestration code honestly. The bugs were all in the unhappy paths β€” fallback, cancellation, failure cleanup β€” exactly the paths I had tested least.


πŸ“š What I Learned

I learned that unhappy paths are where automation rots. Everything works in the demo; the fallback, the cancel, and the crash are where wrong trust decisions and stray processes live.


➑ Next Steps

  • Re-run benchmarks after every model change
  • Add checks for orphaned processes to routine maintenance
  • Keep routing decisions written down with their evidence
  • Extend the harness with security-flavored coding tasks

🧠 Reflection

This was useful because it applied the audit mindset to my own tooling: measure first, route on evidence, and treat every unhappy path as a finding.


🧩 Lessons Learned

What worked

Benchmarking before trusting, and reviewing the orchestrator like third-party code.

What broke

Fallback routing, provider defaults, cancellation, and process cleanup.

Why it broke

I had only exercised the happy path during development.

Fix / takeaway

Test the failure paths deliberately; that is where reliability and security overlap.


πŸ“ˆ Skill Progression Context

This supports my cybersecurity progression because evidence-based trust decisions and unhappy-path analysis are the same skills used in code review and incident investigation.


πŸ˜„ TL;DR

Benchmarked the stack; the bugs were hiding in the failure paths.