π Day 139 β Benchmarking a Local LLM Coding Stack: Harness, Routing, and Review Findings
π Topic
Building and benchmarking the local LLM coding stack β a test harness, routing rules, and fixing what code review found.
π― Goal
Know, with evidence, which local model to trust for which coding task β and clean up the reliability bugs review surfaced.
π What I Did
I documented and benchmarked the local coding stack: a harness for running models against real tasks, routing rules for which model handles what, and workflow docs so the setup survives model swaps. Then I ran a review pass over my own orchestration code and fixed what it found: a routing fallback that could pick the wrong provider, unsafe provider defaults, a cancellation path that could race, and orphaned processes left behind after failures. I also bootstrapped a proper operating layer for the workstationβs coding agents.
Main areas covered:
- benchmark harness for local coding models
- routing rules based on measured capability
- workflow documentation for model swaps
- routing fallback and provider default fixes
- cancellation safety and orphan process cleanup
- operating layer bootstrap for coding agents
π Key Cybersecurity Connections
Reliability bugs are security bugs waiting for context: an unsafe fallback is an unintended trust decision, a cancellation race is a state-integrity flaw, and orphaned processes are unaccounted-for execution. Reviewing my own automation like hostile code is defensive practice.
π Investigation Questions
- Which model does each task type actually route to, and why?
- What happens when the preferred provider is unavailable?
- Can a cancelled task leave work half-applied?
- Are there processes running that no active task owns?
- Do the benchmark results justify the routing table?
π¨ Detection Opportunities
Potential monitoring ideas:
- fallback routing activating unexpectedly
- orphaned agent processes accumulating
- tasks cancelled but still producing output
- provider configuration drift from defaults
- benchmark regressions after model updates
Example:
project=local-coding-stack
change_type=orchestration_reliability_fix
risk_area=unintended_execution_paths
triage=audit_fallbacks_cancellation_and_orphans
π§ MITRE ATT&CK Techniques
Possible mappings depending on confirmed behavior:
- T1059 β Command and Scripting Interpreter
- T1057 β Process Discovery
- T1565 β Data Manipulation
πΊ Visual Investigation Diagram
Benchmark harness
β Measured capability
β Routing rules
β Review findings
β Fixed, documented stack
β Challenges
The challenge was reviewing my own orchestration code honestly. The bugs were all in the unhappy paths β fallback, cancellation, failure cleanup β exactly the paths I had tested least.
π What I Learned
I learned that unhappy paths are where automation rots. Everything works in the demo; the fallback, the cancel, and the crash are where wrong trust decisions and stray processes live.
β‘ Next Steps
- Re-run benchmarks after every model change
- Add checks for orphaned processes to routine maintenance
- Keep routing decisions written down with their evidence
- Extend the harness with security-flavored coding tasks
π§ Reflection
This was useful because it applied the audit mindset to my own tooling: measure first, route on evidence, and treat every unhappy path as a finding.
π§© Lessons Learned
What worked
Benchmarking before trusting, and reviewing the orchestrator like third-party code.
What broke
Fallback routing, provider defaults, cancellation, and process cleanup.
Why it broke
I had only exercised the happy path during development.
Fix / takeaway
Test the failure paths deliberately; that is where reliability and security overlap.
π Skill Progression Context
This supports my cybersecurity progression because evidence-based trust decisions and unhappy-path analysis are the same skills used in code review and incident investigation.
π TL;DR
Benchmarked the stack; the bugs were hiding in the failure paths.
