πŸ”„ Topic

I worked on an AI Inference Orchestrator: a layer that classifies tasks, chooses candidate models, runs them through a controlled pipeline, retries when appropriate, records cost/savings information, and surfaces routing information on a dashboard.


🎯 Goal

Move from β€œpick a model manually” to a more operational system: classify the task, route to a suitable local model, retry or escalate when needed, and record enough evidence to review the decision later.


πŸ›  What I Did

The orchestrator became more than an advisory router.

Main areas covered:

  • added classify/route/execute/retry/escalate pipeline behavior
  • added a cost and savings ledger
  • added multi-model pipeline support
  • surfaced capability scores and cost data on the existing dashboard
  • fixed a small classifier gap where the literal word security was not in the security keyword bucket
  • auto-discovered live Ollama models and benchmarked new candidates
  • deliberately skipped embedding models because the coding/reasoning benchmark suite did not apply to them
  • kept paid-API escalation as a recommendation instead of wiring it automatically

πŸ”— Key Cybersecurity Connections

This is reliability engineering for AI-assisted security work. If an agent can choose tools and models, then its decisions need to be inspectable. Routing is not just convenience; it affects cost, privacy, output quality, and whether a task gets handled by an appropriate capability.

The classifier gap was small but meaningful. A task saying β€œsecurity issue” should not fall into a generic review bucket only because the keyword list contains related terms but not the obvious word.


πŸ” Investigation Questions

  • How does the system decide which model should handle a task?
  • Can that routing decision be explained after the fact?
  • What happens when the first model fails?
  • Is cost tracked, or invisible?
  • Are benchmark results connected to routing decisions?
  • Are task categories too brittle because of naive keyword matching?

🚨 Detection Opportunities

Monitoring ideas for an AI routing layer:

  • task category unexpectedly changes after a classifier edit
  • model selected despite stale or missing benchmark data
  • retry loop repeats the same failing route without variation
  • cost ledger missing entries for completed tasks
  • dashboard status diverges from router state

Example:

project=ai-inference-orchestrator
signal=security_task_routed_to_generic_review
risk_area=classifier_misrouting
triage=inspect_task_classifier_keyword_bucket_and_route_explanation

🧭 MITRE ATT&CK Techniques

No direct ATT&CK mapping claimed. This work supports defensive automation governance rather than modeling adversary behavior.


πŸ—Ί Visual Investigation Diagram

User task
    ↓
Classifier
    ↓
Candidate model scoring
    ↓
Execute
    ↓
Validate / retry / vary route
    ↓
Record cost + result
    ↓
Surface status on dashboard

⚠ Challenges

The most interesting risk was overbuilding. A simple keyword fix was safe; a full classifier rewrite would have had a bigger blast radius and needed a design decision. I kept the safe fix and documented the deeper limitation instead of gambling.


πŸ“š What I Learned

I learned that routing systems need evidence trails. If I cannot explain why a task went to a model, then routing becomes another opaque automation layer.


➑ Next Steps

  • Decide on a more robust classifier approach before rewriting the keyword logic
  • Keep benchmark ingestion tied to live models, not manually curated guesses
  • Avoid automatic paid-API escalation until the governance rules are explicit
  • Continue surfacing routing decisions in operator-facing dashboards

🧠 Reflection

This felt like moving from β€œI have models” to β€œI have an operating system for models.” The difference is not glamorous; it is routing, logs, costs, tests, and restraint.


🧩 Lessons Learned

What worked

Building a pipeline that records routing, execution, retries, and cost.

What needed restraint

Not rewriting the classifier beyond the safe, obvious fix.

Why it mattered

A routing bug can quietly send work to the wrong capability.

Fix / takeaway

AI routing needs the same review mindset as any other decision engine.


πŸ“ˆ Skill Progression Context

This supports my cybersecurity progression because it practices governance over automation: classification, routing, logging, cost awareness, and controlled escalation.


πŸ˜„ TL;DR

Built a real inference orchestrator layer: classify tasks, choose models, retry intelligently, track costs, and make routing decisions visible.