Adversarial Red-Team
Regression Suite
A red-team suite is only useful if it runs on every change, not once at launch. This battery replays 140/140 adversarial attempts, synthetic corpus, mapped to the OWASP LLM Top 10; 0 get through. The attack-success-rate is the number that must stay at zero for a change to ship.
Attack-Success-Rate, Not "We Tested It"
"We ran some jailbreak prompts" is not a security posture. The measurable version is attack-success-rate (ASR): of a fixed battery of adversarial attempts, how many produced the disallowed behavior. Here the battery is 140/140 caught, synthetic corpus, and the ASR holds at zero. The value of the metric is that it is comparable across every code change, a regression shows up as an ASR above zero.
Attack Classes Covered
A Regression Gate, Not a One-Time Pass
The battery is wired into the release gate: it re-runs on every change and blocks the ship if a single attempt gets through. A model or prompt edit that quietly re-opens an injection path does not reach production, because the ASR spikes above zero and the gate goes RED. Robustness stops being a launch-day claim and becomes a property the pipeline enforces on every commit.
Honest Scope
This is a zero-ASR result on a fixed eval corpus, not an all-time guarantee against every attack that will ever exist. A novel attack outside the battery is exactly what the suite does not yet cover, which is why new attack patterns get added when they appear, and the battery grows. The claim is precise: on this corpus, every attempt is caught, and the gate proves it on every run.
Why This Matters For An AI-Security Team
A security team owning an LLM product needs adversarial coverage that does not decay between audits. A red-team battery bound to the release gate turns a point-in-time pentest into a continuous control: every OWASP LLM class has assertions, every change re-runs them, and the ASR is the single number a reviewer reads to know the posture held. It is the difference between "we tested for injection once" and "injection cannot regress without failing the build."
AI-security discipline · Remote · open to mid-level and senior IC roles
