π Day 208 β Adding a Live Trial Because Passing Fixture Tests Did Not Prove a Local Model Was Ready
π Topic
Adding an explicit, opt-in live-evidence path because deterministic shadow tests cannot prove how a real local model will behave.
π― Goal
Keep fast, repeatable shadow trials while making sure they cannot be mistaken for proof that a real Ollama worker is ready for cutover.
π What I Did
I added a separate live-cutover trial runner for the local AI-OS router. It invokes the normal executor and validator against a real local Ollama model, writes only to a unique disposable workspace path, and appends an auditable result to the existing trial ledger.
It is deliberately narrow:
- real-worker execution is explicitly opt-in;
- each category runs separately, so one honest failure stops that category;
- remote bridges, paid providers, and personal-data surfaces stay disabled;
- the result must include a clean terminal status, return code, and validator result;
- ledger records are labeled
live_opt_in, distinct from deterministic shadow evidence.
The important design choice is separation. The fast test path remains useful, but it cannot accidentally satisfy a decision that requires live evidence.
π Key Cybersecurity Connections
This is evidence integrity and change-management discipline. A controlled test double can prove that a route and validator agree; it cannot prove the live model, runtime, prompt budget, or local environment will behave the same way.
π Investigation Questions
- Was this result produced by a fixture or a real worker?
- Which model and provider produced it?
- Did the validator inspect the expected artifact?
- What paths and external surfaces were in scope?
- Can a stale or simulated result unlock a live decision?
π¨ Detection Opportunities
Track and review:
- cutover decisions with no
live_opt_inledger evidence - model unavailable during preflight
- validation pass with a missing or wrong workspace artifact
- attempted use of a remote bridge during a local-only trial
Example:
event=cutover_evidence_rejected
reason=shadow_result_not_live_opt_in
action=require_scoped_real_worker_trial
π§ MITRE ATT&CK Techniques
There is no direct ATT&CK technique here; this is defensive assurance work. It supports reliable analysis by preventing false confidence in automation controls.
πΊ Visual Investigation Diagram
Deterministic shadow trial βββ useful regression evidence
β
βββ not a cutover decision
Explicit local Ollama trial βββ executor + validator + unique artifact
β
live_opt_in ledger evidence
β Challenges
It is easy to call a test βreal enoughβ when it passes. The harder, more honest move is to say exactly what it did not test and give the missing evidence a different shape.
π What I Learned
I learned that evidence labels are security controls. If fixture evidence and live evidence look interchangeable in a ledger, a future reader can make a high-impact decision from the wrong kind of proof.
β‘ Next Steps
- Run categories one at a time and retain only factual results
- Review the ledger before declaring a route ready
- Keep the live path local, low-risk, and explicitly invoked
π§ Reflection
The goal is not to distrust tests. It is to preserve what each test actually means.
π§© Lessons Learned
What worked
Separate evidence modes and require a complete result contract.
What broke
The assumption that a green simulated worker said enough about a live one.
Why it broke
Simulation removes exactly the runtime uncertainty a cutover needs to address.
Fix / takeaway
Use shadow trials for regression confidence and scoped live trials for live-readiness evidence.
π Skill Progression Context
This strengthened my understanding that validation is not a single checkbox. The source, scope, and realism of evidence matter.
π TL;DR
A passing fixture is valuableβbut it is not proof that a real local model is ready to take over.
