Adversarial LLM Defense Framework
Public Adversarial Gate (5 OWASP classes, 140/140) + Multi-Agent Council (production, NDA)
The Problem
Production LLM systems in regulated industries (banking, healthcare, defense) have no deterministic audit trail. SR 11-7 (Federal Reserve model risk guidance) requires explainability and validation that LLM vendors do not provide out of the box. Standard SAST/DAST tooling was not built for prompt injection, jailbreak, or tool-abuse classes — the OWASP LLM Top 10 sits outside its coverage.
Approach
Public gate (clone & run): 5 OWASP LLM Top 10 threat classes, 140/140 adversarial caught, 0/80 benign over-blocked, 114 offline pytest assertions on a bundled n=220 corpus. The production deployment extends this — multi-agent adversarial validation council, SHA-256 hash-chained replay corpus, SR 11-7 alignment — under NDA. Methodology public via the gate repo.
Proof Metrics
Want this on your team?
This work maps to 5 qualifying roles. Remote, Q3 2026.