π Day 151 β Benchmarking My Local LLM Stack Instead of Trusting Vibes
π Topic
I compared the local model stack on my laptop and made a concrete routing decision: Ollama with my juribuora-agent-coder:30b model is the daily backend, DS4 stays installed but experimental, and LM Studio remains useful for GUI experiments but unreliable for headless automation in this setup.
π― Goal
Stop choosing local models by hype, size, or vibes. Measure which backend is fast, coherent, stable, and safe enough for agentic coding workflows.
π What I Did
I reviewed and verified the local AI stack around Ollama, DS4, LM Studio, Aider, Codex CLI, and my local-agent harness.
Main areas covered:
- benchmarked DS4 with
deepseek-v4-flash - benchmarked Ollama models including
juribuora-agent-coder:30bandqwen3-coder:30b - checked LM Studioβs headless API behavior
- confirmed Aider and the local-agent harness could run against the local stack
- documented routing rules for when to use local models and when to escalate
- fixed a health-check script bug where
rgwas assumed to exist on the PATH - tested a real local-agent dry run and confirmed its write boundaries are intentionally restricted
π Key Cybersecurity Connections
This connects to security in two ways. First, local models reduce exposure of sensitive project context because more work can stay on the machine. Second, autonomy needs evidence. If a model is slow, unstable, or returns garbage through an API path, it should not be trusted just because it looks impressive in a GUI.
The write-scope boundary was also important: my local-agent cannot edit arbitrary source files. That is a deliberate containment decision, not a missing feature.
π Investigation Questions
- Which model is actually fastest on this hardware?
- Which model produces useful coding output consistently?
- Does the GUI behavior match the headless API behavior?
- Can the agent harness write only where it is supposed to write?
- When should a local model be used, and when should the task escalate to a stronger external model?
π¨ Detection Opportunities
Useful checks for the local stack:
- local model health check fails because a tool is missing from PATH
- headless API returns incoherent output even though the GUI looks healthy
- task asks local-agent to write outside approved directories
- same task fails twice locally and should be escalated
- benchmark results drift after a model, runtime, or configuration change
Example:
project=local-llm-stack
signal=headless_api_output_invalid
risk_area=automation_model_reliability
triage=compare_gui_output_api_output_and_benchmark_logs
π§ MITRE ATT&CK Techniques
No direct mapping claimed. This is infrastructure reliability and data-exposure reduction for my own tooling.
πΊ Visual Investigation Diagram
Candidate local models
β
Benchmark speed and quality
β
Test headless API behavior
β
Verify tool harness boundaries
β
Choose default route
β
Escalate only when local attempts fail or review is needed
β Challenges
The hardest part was not the benchmark itself. It was resisting the instinct to use the biggest or newest model by default. DS4 was impressive but too heavy as an always-on coding backend. LM Studio was promising in the GUI but not trustworthy through the tested API path.
π What I Learned
I learned that model selection is an operations decision, not a fan-club decision. The best default is the model that repeatedly performs the actual job under the actual constraints.
β‘ Next Steps
- Keep Ollama
juribuora-agent-coder:30bas the daily default - Use DS4 only for explicit heavy-reasoning experiments
- Treat LM Studio headless automation as untrusted until the output issue is fixed
- Re-run benchmarks after model/runtime updates
- Keep local-agent write boundaries documented and enforced
π§ Reflection
This was one of those days where βboring measurementβ beat excitement. The biggest model was not the most useful default. The reliable model was.
π§© Lessons Learned
What worked
Benchmarking the same practical task across runtimes instead of guessing.
What broke
LM Studioβs tested headless API path produced unusable output.
Why it broke
The GUI being alive did not prove the API path was suitable for automation.
Fix / takeaway
Every automation backend needs its own verification path. Pretty UI output does not certify headless agent behavior.
π Skill Progression Context
This supports my cybersecurity progression because it builds disciplined evidence-based decision-making: measure the control, verify the boundary, document the fallback.
π TL;DR
Benchmarked the local model stack and chose Ollama as the daily backend, kept DS4 experimental, and learned again that vibes are not a routing policy.
