🔄 Topic

Rolling out a WhatsApp-facing persona for Hermes meant I could no longer eyeball a handful of test messages and call it safe. I built an automated adversarial harness to attack the persona itself — and then had to fix the harness twice before I trusted what it was telling me.


🎯 Goal

Get a repeatable, automated way to measure whether the persona holds up under adversarial pressure — factual questions, capability probing, language edge cases — instead of relying on a spot-check that only proves the happy path works.


🛠 What I Did

I built the harness, then treated its own failures as bugs to fix rather than results to accept at face value.

Main areas covered:

  • built an automated adversarial harness for the WhatsApp persona: scripted adversarial prompts run against it repeatedly, rather than manual one-off testing before a rollout
  • added --repeat to the harness to measure flake rate directly — running the same adversarial prompt multiple times and checking whether the persona’s behavior is consistent, not just correct once
  • fixed a language-rule bug caught by the repeated runs along the way
  • found the harness itself was failing correct answers — a false-positive in the test tooling, not the persona — and fixed the harness rather than loosening the assertions to make the failures disappear
  • extended the harness to assert facts and specifically catch feminine grammatical agreement errors (a real, checkable failure mode for a persona operating in Italian), and to “ask the real filter” — testing the actual decision path rather than a proxy for it
  • changed how capability disclosure is scored: instead of a binary pass/fail on whether the persona revealed internal capabilities when probed, the harness now measures and records disclosure behavior, since the honest answer to “does this system talk about what it can do” needed data, not a single verdict
  • fixed a related generator bug: persona rules were being owned in the wrong place, and there was a midnight gap where date-dependent logic could misbehave right at the day boundary — both surfaced by the same harness runs

🔗 Key Cybersecurity Connections

Adversarial testing of a conversational AI system is the same discipline as fuzzing or penetration testing applied to a language interface: you don’t trust that a system is safe because it behaved correctly the one time you tried it, you actively try to break it, repeatedly, and measure the failure rate rather than treating a single pass as proof. Flake-rate measurement specifically matters for LLM-backed systems because non-determinism means “it worked when I tested it” is a much weaker claim than for traditional deterministic code — a prompt that passes once and fails one time in five is not a passing prompt.

Trusting the test tooling itself is its own security discipline: a harness that fails correct answers will either get “fixed” by loosening assertions (masking real regressions later) or ignored (defeating the point of having it). Fixing the harness’s own bugs, rather than papering over its false positives, is what makes its future results trustworthy.


🔍 Investigation Questions

  • Does the persona’s behavior stay consistent across repeated runs of the same adversarial prompt, or does it flake?
  • When the harness reports a failure, is it a real persona bug or a bug in the harness’s own assertions?
  • Does the persona correctly disclose its capabilities when directly asked, and is that measured rather than assumed?
  • Are language-specific failure modes (like grammatical agreement in a non-English persona) actually tested, or only English-shaped bugs?
  • Is a rule generator’s ownership boundary clear, or can the same logic drift into being defined in two places?

🚨 Detection Opportunities

Checks for an adversarial test harness on a conversational system:

  • a persona change shipped without a corresponding adversarial harness run
  • a harness assertion loosened to make a failure disappear, instead of the underlying bug fixed
  • no flake-rate measurement — correctness checked once instead of across repeated runs
  • capability-disclosure behavior unmeasured, relying on assumption rather than a scored test
  • a date- or time-boundary condition (like a midnight gap) untested by the harness

Example:

project=whatsapp-persona-harness
signal=harness_assertion_loosened_instead_of_bug_fixed
risk_area=false_confidence_in_adversarial_test_coverage
triage=revert_assertion_change_confirm_underlying_persona_behavior_first

🧭 MITRE ATT&CK Techniques

No direct mapping claimed. This is adversarial/red-team testing methodology applied to an LLM-backed conversational system, adjacent to AI-specific frameworks like MITRE ATLAS rather than classic ATT&CK.


🗺 Visual Investigation Diagram

WhatsApp persona rollout planned
    ↓
Manual spot-checking isn't enough for repeatable confidence
    ↓
Build automated adversarial harness
    ↓
Add --repeat: measure flake rate, not just single-pass correctness
    ↓
Harness itself found failing correct answers
    ↓
Fix the harness, not the assertions
    ↓
Add fact assertions, language-agreement checks, real capability-disclosure scoring
    ↓
Trustworthy, repeatable adversarial signal before rollout

⚠ Challenges

The hardest part was resisting the urge to treat every harness failure as a persona bug. Twice the actual bug was in the harness — and it would have been easy, and wrong, to “fix” those by loosening what the harness checked for, which would have quietly reduced coverage while looking like progress.


📚 What I Learned

I learned that a test harness is itself untrusted code until proven otherwise, and that “the harness says it failed” and “the system actually failed” are two different claims that both need verifying. I also learned that scoring capability disclosure as measured behavior, rather than a binary pass/fail, gave me a more honest picture than forcing a verdict onto something that’s genuinely a spectrum.


➡ Next Steps

  • Run the harness on a schedule against the live persona, not just before a rollout
  • Extend flake-rate measurement to other language-dependent behaviors beyond grammatical agreement
  • Track capability-disclosure scores over time to catch drift, not just a one-time baseline
  • Apply the same harness pattern to other persona-bearing channels as they’re added

🧠 Reflection

Building a tool whose entire job is to try to break something I built is a different mindset than building the thing itself, and it’s the mindset that actually earns confidence — a persona that survives repeated, hostile, automated pressure is trusted for a real reason, not because nobody tried hard enough to break it.


🧩 Lessons Learned

What worked

Measuring flake rate across repeated adversarial runs instead of trusting a single pass, and fixing harness bugs instead of loosening assertions.

What broke

The harness itself twice — failing correct answers — plus a real language-agreement bug and a persona-rule generator ownership/midnight-gap issue it caught.

Why it mattered

A conversational AI persona is non-deterministic, so single-run testing is a much weaker guarantee than for deterministic code, and an untrustworthy test harness defeats the entire point of automated adversarial coverage.

Fix / takeaway

Treat the test harness as code that can itself be wrong, measure consistency across repeats rather than one-off correctness, and score nuanced behaviors like capability disclosure rather than forcing a binary verdict.


📈 Skill Progression Context

This supports my cybersecurity progression because adversarial testing methodology, flake-rate measurement, and distrust of your own test tooling until verified are directly transferable red-team and QA-security skills, applied here to an LLM-backed system instead of a traditional application.


😄 TL;DR

Built an automated harness to attack my own chatbot’s persona — then had to fix the attacker before I could trust what it was telling me about the target.