Skip to main content
puodziukas.dev›CASES/DETERMINISTIC GATE BATTERY/
CASE STUDY · VERIFICATION ENGINEERING · QA INFRASTRUCTURE

Deterministic Gate Battery:
Model Output as Claim

A deterministic gate battery · 2 false-green AI reports caught
An AI model reported a text contrast ratio of 5.52:1 on both the light theme AND the dark theme — the same ratio on two different backgrounds. Impossible. Recomputed from relative luminance: the dark theme was 3.05:1, an accessibility failure labeled as a pass. A second model regressed a 10.16:1 score and reported it as improvement. Neither shipped.

The Failures

AI model #1 (accessibility audit)
AI reported: Light-theme contrast ratio: 5.52:1. Dark-theme contrast ratio: 5.52:1.
Reality: Recomputed from WCAG AA relative luminance: light 5.52:1 ✓, dark 3.05:1 ✗
Caught by: Gate 7 (contrast-usage-aware): threshold floor 4.5:1, calculates per-theme from pixel samples
No — gate RED, rejected before preview
AI model #2 (design readability pass)
AI reported: Body text accessibility score improved from 10.16:1 to 6.81:1.
Reality: The ratio regressed. 10.16:1 > 6.81:1. Still above AA floor (4.5:1), but direction is reversal.
Caught by: Gate 10 (canon-sync): commits must list the 55 live metrics verbatim; regressed metrics spike the diff
No — gate RED, commit aborted

What Mutation-Proven Means

A gate is only counted after it has flipped RED on a planted defect AND stayed GREEN on a coincidental-green control. This proof standard is higher than "the gate works on real defects" — it requires explicit evidence that the gate actually detects the failure class it claims to cover.

The gates in the battery run on every commit. Two of them caught AI model runs that reported a pass on work that had failed. The gate battery is not a false-positive-free system, it is a "catches what it says it catches" system.

The Gate Battery

A deterministic gate battery runs on every ship; a representative sample is shown below.

GATE 01
TypeScript: zero implicit any
GATE 02
Build: tsc --noEmit exits 0
GATE 03
Palette: royal blue + neutrals only (no hex literals)
GATE 04
Mobile: 375px viewport passes all interactions
GATE 05
Semantics: no div when button needed
GATE 06
Contrast: WCAG AA ≥4.5:1 pre-verified per theme
GATE 07
Contrast-usage-aware: scores mutations + regression detection
GATE 08
Accessibility: ARIA, skip-links, 44px touch targets
GATE 09
Copy lexicon: build fails on any of 30+ locked marketing-filler terms
GATE 10
Canon sync: 55 reproducible metrics must match anchors.ts + claims-canon.json
GATE 11
Claims gate: every claim has a public replay path, or it is cut
GATE 12
Routes: no broken links on live server
GATE 13
Performance: LCP gate, CLS gate, JS bundle size budget
GATE 14
SEO: unique title+desc+canonical per route
GATE 15
Security: zero secrets, full HSTS/CSP headers, npm audit 0 high/crit

Why This Matters

In a live system, an AI model is an oracle, not a fact source. When a model reports "accessibility passed," the statement is data, not truth. It is a claim that requires independent verification before it reaches a decision boundary.

A deterministic gate battery is the layer that translates claims into provable statements. The gate knows the implementation detail (pixel colors, WCAG definitions, live file state). The model does not. The architecture that ships is: model generates output → human designs a deterministic oracle that checks the output → the oracle's verdict is what ships.

Two AI reports were rejected by this battery. The same battery would reject thousand more, because it is not tuned to the 2 reports — it is tuned to the failure classes those 2 reports exemplify: scoring systems that do not understand context, and changes that invert their own metrics.

# Verification engineering contract if model_output_passes: # model_output is a CLAIM for gate in gate_battery: gate_verdict = oracle(model_claim, ground_truth) if gate_verdict == RED: # reject, do not ship return FAIL # all gates GREEN: consensus between model and deterministic oracle return SHIP
PROOF
✓Deterministic gate battery - Every check; one RED blocks the commit
✓2 false-greens caught - Contrast ratio impossibility + score regression reported as improvement
✓0 bypasses - No SKIP_GATE flags, no emergency overrides, no exceptions logged
✓Reproducible - Gate definitions locked in code, every check is deterministic and replayable
Work with me →

verification-proven · Remote · open to mid-level and senior IC roles

RELATED CASES
Eval & Release Gate →Agentic Fleet + Deterministic Verification →