πŸ”„ Topic

Adding an explicit, opt-in live-evidence path because deterministic shadow tests cannot prove how a real local model will behave.


🎯 Goal

Keep fast, repeatable shadow trials while making sure they cannot be mistaken for proof that a real Ollama worker is ready for cutover.


πŸ›  What I Did

I added a separate live-cutover trial runner for the local AI-OS router. It invokes the normal executor and validator against a real local Ollama model, writes only to a unique disposable workspace path, and appends an auditable result to the existing trial ledger.

It is deliberately narrow:

  • real-worker execution is explicitly opt-in;
  • each category runs separately, so one honest failure stops that category;
  • remote bridges, paid providers, and personal-data surfaces stay disabled;
  • the result must include a clean terminal status, return code, and validator result;
  • ledger records are labeled live_opt_in, distinct from deterministic shadow evidence.

The important design choice is separation. The fast test path remains useful, but it cannot accidentally satisfy a decision that requires live evidence.


πŸ”— Key Cybersecurity Connections

This is evidence integrity and change-management discipline. A controlled test double can prove that a route and validator agree; it cannot prove the live model, runtime, prompt budget, or local environment will behave the same way.


πŸ” Investigation Questions

  • Was this result produced by a fixture or a real worker?
  • Which model and provider produced it?
  • Did the validator inspect the expected artifact?
  • What paths and external surfaces were in scope?
  • Can a stale or simulated result unlock a live decision?

🚨 Detection Opportunities

Track and review:

  • cutover decisions with no live_opt_in ledger evidence
  • model unavailable during preflight
  • validation pass with a missing or wrong workspace artifact
  • attempted use of a remote bridge during a local-only trial

Example:

event=cutover_evidence_rejected
reason=shadow_result_not_live_opt_in
action=require_scoped_real_worker_trial

🧭 MITRE ATT&CK Techniques

There is no direct ATT&CK technique here; this is defensive assurance work. It supports reliable analysis by preventing false confidence in automation controls.


πŸ—Ί Visual Investigation Diagram

Deterministic shadow trial ──→ useful regression evidence
                                  β”‚
                                  └── not a cutover decision

Explicit local Ollama trial ──→ executor + validator + unique artifact
                                  ↓
                          live_opt_in ledger evidence

⚠ Challenges

It is easy to call a test β€œreal enough” when it passes. The harder, more honest move is to say exactly what it did not test and give the missing evidence a different shape.


πŸ“š What I Learned

I learned that evidence labels are security controls. If fixture evidence and live evidence look interchangeable in a ledger, a future reader can make a high-impact decision from the wrong kind of proof.


➑ Next Steps

  • Run categories one at a time and retain only factual results
  • Review the ledger before declaring a route ready
  • Keep the live path local, low-risk, and explicitly invoked

🧠 Reflection

The goal is not to distrust tests. It is to preserve what each test actually means.


🧩 Lessons Learned

What worked

Separate evidence modes and require a complete result contract.

What broke

The assumption that a green simulated worker said enough about a live one.

Why it broke

Simulation removes exactly the runtime uncertainty a cutover needs to address.

Fix / takeaway

Use shadow trials for regression confidence and scoped live trials for live-readiness evidence.


πŸ“ˆ Skill Progression Context

This strengthened my understanding that validation is not a single checkbox. The source, scope, and realism of evidence matter.


πŸ˜„ TL;DR

A passing fixture is valuableβ€”but it is not proof that a real local model is ready to take over.