π Day 212 β Choosing a Local Coding Model by Measuring Repeated Full Rounds
π Topic
Choosing a local coding model by measuring complete repeated rounds, not by model size or a one-off first response.
π― Goal
Make the agent bridge responsive without overloading the laptop or trusting a model that appears available but cannot reliably complete a round.
π What I Did
I benchmarked local worker candidates against the same roughly 10.5K-token prompt on a quiet machine. The important result was not simply which model had the fastest cold response. It was whether the repeated prompt prefix was actually cached.
The measured result favored the GPT-OSS 120B worker: it produced a usable answer in 17.5 seconds cold and 2.2 seconds warm. Other candidates had respectable first calls but no useful warm-cache improvement, or returned empty output after consuming their token budget.
I also reduced the workerβs inherited context by removing tool and skill material it never used. That cut a representative round from 79 seconds to 26 seconds. A separate protection keeps only one heavy model resident when the machine is busy, instead of pretending multiple GPU-hungry models can coexist indefinitely.
π Key Cybersecurity Connections
Performance is part of availability. A local security workflow that times out, swaps heavily, or returns empty output during a real task is not operationally reliable just because its endpoint is listening.
π Investigation Questions
- Is a model returning usable content, not only a successful HTTP response?
- Does the repeated system prompt hit a real cache?
- Was the benchmark run on a quiet machine?
- How much context is permanent overhead on every turn?
- Which model is actually resident on the GPU?
π¨ Detection Opportunities
Useful telemetry:
- cold versus warm latency
- prompt-cache reuse
- empty-content completions
- token-limit finishes
- swap pressure and concurrent model residency
Example:
event=local_worker_round
model=gpt-oss-120b
warm_latency_seconds=2.2
output=usable
action=retain_as_bridge_worker
π§ MITRE ATT&CK Techniques
This is resilience engineering rather than a direct ATT&CK mapping. It supports reliable defensive automation under realistic resource constraints.
πΊ Visual Investigation Diagram
Same prompt, quiet machine
β
Cold response + warm response + usable output
β
Cache behavior + context cost + residency pressure
β
Choose a route based on evidence
β Challenges
An early measurement looked worse after a configuration change, but it was polluted by cache invalidation and GPU contention. Without controlling those variables, it would have led to the wrong model decision.
π What I Learned
I learned that repeated-workload performance matters more than a headline speed number. For an agent loop, every unnecessary prompt token and every cache miss becomes a tax paid again and again.
β‘ Next Steps
- Re-benchmark after meaningful model or prompt changes
- Keep the worker context scoped to the work it actually performs
- Report unavailable or empty-output models honestly
π§ Reflection
The model that βlooks loadedβ is not necessarily the model that can finish work. Availability needs evidence from the whole round.
π§© Lessons Learned
What worked
Measuring warm and cold behavior, output quality, and resource contention together.
What broke
Choosing from a single timing result on a busy machine.
Why it broke
The GPU, cache, context window, and concurrent jobs all changed the result.
Fix / takeaway
Benchmark the real repeated workflow, on a controlled machine state.
π Skill Progression Context
This is the operational side of local AI security work: capacity, latency, and truthful monitoring determine whether a control is usable when needed.
π TL;DR
The winning local worker was not merely quickβit was the one that reused the prompt and returned usable work reliably.
