CASE STUDY · VERIFICATION ENGINEERING · QA INFRASTRUCTURE
Deterministic Gate Battery: Model Output as Claim
A deterministic gate battery · 2 false-green AI reports caught
An AI model reported a text contrast ratio of 5.52:1 on both the light theme AND the dark theme — the same ratio on two different backgrounds. Impossible. Recomputed from relative luminance: the dark theme was 3.05:1, an accessibility failure labeled as a pass. A second model regressed a 10.16:1 score and reported it as improvement. Neither shipped.
The Failures
AI model #1 (accessibility audit)
AI reported: Light-theme contrast ratio: 5.52:1. Dark-theme contrast ratio: 5.52:1.
Reality: Recomputed from WCAG AA relative luminance: light 5.52:1 ✓, dark 3.05:1 ✗
AI reported: Body text accessibility score improved from 10.16:1 to 6.81:1.
Reality: The ratio regressed. 10.16:1 > 6.81:1. Still above AA floor (4.5:1), but direction is reversal.
Caught by: Gate 10 (canon-sync): commits must list the 55 live metrics verbatim; regressed metrics spike the diff
No — gate RED, commit aborted
What Mutation-Proven Means
A gate is only counted after it has flipped RED on a planted defect AND stayed GREEN on a coincidental-green control. This proof standard is higher than "the gate works on real defects" — it requires explicit evidence that the gate actually detects the failure class it claims to cover.
The gates in the battery run on every commit. Two of them caught AI model runs that reported a pass on work that had failed. The gate battery is not a false-positive-free system, it is a "catches what it says it catches" system.
The Gate Battery
A deterministic gate battery runs on every ship; a representative sample is shown below.
GATE 01
TypeScript: zero implicit any
GATE 02
Build: tsc --noEmit exits 0
GATE 03
Palette: royal blue + neutrals only (no hex literals)
Security: zero secrets, full HSTS/CSP headers, npm audit 0 high/crit
Why This Matters
In a live system, an AI model is an oracle, not a fact source. When a model reports "accessibility passed," the statement is data, not truth. It is a claim that requires independent verification before it reaches a decision boundary.
A deterministic gate battery is the layer that translates claims into provable statements. The gate knows the implementation detail (pixel colors, WCAG definitions, live file state). The model does not. The architecture that ships is: model generates output → human designs a deterministic oracle that checks the output → the oracle's verdict is what ships.
Two AI reports were rejected by this battery. The same battery would reject thousand more, because it is not tuned to the 2 reports — it is tuned to the failure classes those 2 reports exemplify: scoring systems that do not understand context, and changes that invert their own metrics.
# Verification engineering contract
if model_output_passes:
# model_output is a CLAIM
for gate in gate_battery:
gate_verdict = oracle(model_claim, ground_truth)
if gate_verdict == RED:
# reject, do not ship
return FAIL
# all gates GREEN: consensus between model and deterministic oracle
return SHIP
PROOF
✓Deterministic gate battery - Every check; one RED blocks the commit
✓2 false-greens caught - Contrast ratio impossibility + score regression reported as improvement
✓0 bypasses - No SKIP_GATE flags, no emergency overrides, no exceptions logged
✓Reproducible - Gate definitions locked in code, every check is deterministic and replayable