EVIDENCE · SEVEN ARTIFACTS · PUBLIC ONES CLONE-AND-RUN
Seven evidence artifacts. The public ones clone and run; the rest are labeled, not dressed up.
The public artifacts (cobol-pic-probe, llm-adversarial-gate) are independently replayable. Clone and run. Private and in-progress work is labeled as such, never dressed up as a public replay; where a corpus ships, it is synthetic and fixed-seed.
LIVE METHODOLOGY · MULTI-AGENT COUNCIL · COBOL AUDIT TRACE
The COBOL PICTURE-clause forensic audit was a Reflexion loop: the adversarial agent blocked the scope, the critic blocked the analyst, the truncation defect was caught on the second pass. The trace below is real. Not a simulation.
One instruction given. A multi-agent deliberative council ran the scope: one agent blocked another when it was incomplete, the second pass found additional risk, and the final decision was SHA-256 hash-chained before any work began. Below is the full trace: every agent turn, every block, every replan. Click any row to expand.
council-replay --session 2026-0525-0914 --hash 3a7fc291Work with me →SEVEN ARTIFACTS · WHY THE COMBINATION IS RARE
Consultants produce reports. This record is reproducible where it's public: every public claim ships a deterministic replay command (clone and run); the private and in-progress ones are labeled, not dressed up.
- COBOL PICTURE-clause fixed-point truncation forensic audit: failure class compounding undetected below totals-reconciliation thresholds. Demonstrated on a synthetic, fixed-seed corpus, $40,812.81 compounding loss across the replay corpus (sub-cent per record, 0.0900% of total), 31 pytest assertions. Clone and run: github.com/mpuodziukas-labs/cobol-pic-probe (CI-verified, deterministic, no client data).
- LLM adversarial-validation gate: 5 OWASP LLM Top 10 threat classes, full adversarial block rate (breakdown in the proof log below), 0% false-positive (0/80 benign), 114 pytest assertions on a bundled, offline corpus (n=220). Clone, extend, run: github.com/mpuodziukas-labs/llm-adversarial-gate.
- 158 TPS aggregate · 389ms P50 end-to-end · 531ms TTFT · $0 API cost: sovereign cluster, from $26,400/mo API. Self-hosted — no managed-API dependency, hardware/power cost still applies. SHA-256 hash-chained outputs.
- COBOL→Rust throughput: full benchmark in the proof log below — 464 tests, zero known CVEs.
- Post-quantum migration readiness (FIPS 203/204/205): the three final NIST PQC standards and the NIST 2030 migration timeline mapped; HSM-bridge architecture and harvest-now-decrypt-later exposure assessed. Advisory, not a deployed implementation.
- SHA-256 hash-chained output integrity: every inference result deterministically replayable to its exact input. Clone and run.
- Deterministic audit trail, hash-chained: QMS, bias, and human-oversight documentation built and SHA-256 hash-chained end-to-end — engineering artifacts, not a legal opinion.
Seven numbers. The public ones reproducible from the artifact they ship with; the rest labeled honestly.
On working methodThe replay archive is built alongside the work, not assembled after the fact when someone asks for it. That is what keeps every claim on this page reproducible rather than asserted.
WRITTEN SCOPE FIRST · FIRST ARTIFACT IN 48H · REPLAY ON EVERY CLAIM
The replay command ships with the work.
TRY THE METHODOLOGY · ONE SMALL TOOL, OPEN TO ANYONE
Run it yourself.
COBOL arithmetic risk scan — the same tool used in the work above, reproducible, with a replay command on every result.
COBOL risk pattern scanner
12 patterns sourced to FFIEC IT Handbook + OCC Bulletin 2014-13 + SOX §404 control gaps. Paste code below.
Methodology not on this page. The replay command reproduces in a technical screen. Everything below is independently verifiable.
5 of 5 deterministic build gates passed at this deploy. Same gates, every build, no exceptions.
Happy to walk through the architecture in a technical screen.
Work with me →Q3 2026 · 1 slot