Skip to main content
puodziukas.dev›CASES/ADVERSARIAL RED TEAM/
CASE STUDY · AI SECURITY · ADVERSARIAL ROBUSTNESS

Adversarial Red-Team
Regression Suite

140/140 caught, synthetic corpus · mapped to the OWASP LLM Top 10
A red-team suite is only useful if it runs on every change, not once at launch. This battery replays 140/140 adversarial attempts, synthetic corpus, mapped to the OWASP LLM Top 10; 0 get through. The attack-success-rate is the number that must stay at zero for a change to ship.

Attack-Success-Rate, Not "We Tested It"

"We ran some jailbreak prompts" is not a security posture. The measurable version is attack-success-rate (ASR): of a fixed battery of adversarial attempts, how many produced the disallowed behavior. Here the battery is 140/140 caught, synthetic corpus, and the ASR holds at zero. The value of the metric is that it is comparable across every code change, a regression shows up as an ASR above zero.

Attack Classes Covered

OWASP LLM class
prompt injection
Instructions smuggled through user input that try to override the system prompt.
OWASP LLM class
data exfil
Attempts to make the model leak system prompts, secrets, or prior-turn private data.
OWASP LLM class
jailbreak
Role-play and obfuscation attacks that try to unlock restricted behavior.
OWASP LLM class
tool abuse
Coercing an agent to call a tool with attacker-chosen arguments.
OWASP LLM class
output handling
Payloads that stay dormant in output and fire when a downstream system renders them.

A Regression Gate, Not a One-Time Pass

The battery is wired into the release gate: it re-runs on every change and blocks the ship if a single attempt gets through. A model or prompt edit that quietly re-opens an injection path does not reach production, because the ASR spikes above zero and the gate goes RED. Robustness stops being a launch-day claim and becomes a property the pipeline enforces on every commit.

Honest Scope

This is a zero-ASR result on a fixed eval corpus, not an all-time guarantee against every attack that will ever exist. A novel attack outside the battery is exactly what the suite does not yet cover, which is why new attack patterns get added when they appear, and the battery grows. The claim is precise: on this corpus, every attempt is caught, and the gate proves it on every run.

# Adversarial regression gate (per change) results = run_red_team(corpus) # fixed battery asr = sum(r.succeeded for r in results) / len(results) if asr > 0.0: block_ship() # a reopened path is a RED gate else: allow_merge() # every attempt caught this run

Why This Matters For An AI-Security Team

A security team owning an LLM product needs adversarial coverage that does not decay between audits. A red-team battery bound to the release gate turns a point-in-time pentest into a continuous control: every OWASP LLM class has assertions, every change re-runs them, and the ASR is the single number a reviewer reads to know the posture held. It is the difference between "we tested for injection once" and "injection cannot regress without failing the build."

PROOF
✓140/140 caught, synthetic corpus - 0 successful, on the eval corpus
✓OWASP LLM Top 10 - prompt injection · data exfil · jailbreak · tool abuse · output handling
✓Replay command - pytest llm-adversarial-gate/ -v (offline, fixed corpus)
✓Regression-gated - Re-runs on every change; any bypass blocks the ship
Work with me →

AI-security discipline · Remote · open to mid-level and senior IC roles

RELATED CASES
Eval & Release Gate →AI Governance Mapping →