47,312 REPLAYS · 0 DIVERGENCE · 55 COMMITS AHEAD OF UPSTREAM · COBOL FIND IN 127 MIN
I build the system and red-team it in the same pass.
The failure class surfaces in the build, not in production. Written scope, first artifact in 48 hours. Senior / Staff / Principal · remote.
One operator, end to end — no handoff between a build team and a review team, no status-meeting drag, no second contractor needed to catch what the first missed.
Work with meBefore → After. Every public number has a replay command.
The core claims clone and run from public repos. The cluster reproduces live in a technical screen; NDA evals reproduce on request.
How it works
- 01I work from a written spec — no discovery phase, first artifact fast.
- 02First artifact lands in 48 hours, with its replay command.
- 03Async by default — code lands in your repo, review artifacts not standups. Decisions arrive as written diffs with replay commands; no status calls, no calendar dependency.
- 04Every output is reproducible — clone it and check.
FIRST 30 DAYS
Week 1
First artifact shipped — a deterministic replay command on every LLM decision in the audit trail.
Week 3
Every COBOL arithmetic boundary re-derived. ON SIZE ERROR gaps identified and documented with fix spec.
Day 45
Board question answered with a log line, not a policy document — the replay command in the artifact itself.
Worth a conversation if
- +The team runs an agent fleet, an inference bill, or a legacy system that has to be correct under adversarial conditions — and needs one remote IC who ships replay-verified artifacts, not status updates.
- +The team measures output by artifacts and replay commands, not standups.
- +The role is remote and the scope is written down before day one.
Not a fit if
- –The role needs someone to convince the org the problem is real.
- –The role is management, not building.
- –On-site, daily-standup, ticket-velocity shops.
“What does an internal team not already have?”
structural distance · zero false negatives · three concurrent systems · NDA
A 36-month production run, zero false negatives, across three concurrent systems — the probe arrives with no prior attachment to the codebase. That structural distance is what an internal team can't easily replicate: the team that built the system has a harder time auditing it from outside. It's part of why a truncation defect can age through years of internal audits that were never designed to surface it.
“How do you verify the work without a client list?”
replay ships with every deliverable · clone-verify from input hash
The proof is the artifact. Every deliverable ships with the replay command that reconstructs its result from the input hash — clone it, run it, check it against the documented number. The methodology was validated in production before this became a client-facing offering, so what you're auditing is the log, not a reference call.
Open to one Staff/Principal role, Q3 2026. Three questions: the role, the system, the timeline.
GET IN TOUCH · 24-HOUR REPLY
One paragraph. 24-hour reply.
The more specific the note, the more useful the reply.