AI Evaluation Engineer
Open to mid-level and senior IC roles, W2, remote (US).
I build the systems that run evaluation tests and check that their results are correct.
eval runcontrol PASSplanted regression FAILgatereceipt
Requirements and proof
| Requirement | Proof |
|---|---|
| Build the backend platform that executes evaluation tests and processes their results | Release-gate case study |
| Improve the correctness, performance, and observability of a test-execution system | Release-gate case study - Deterministic gate battery case study |
| Investigate a live issue to root cause and ship a durable fix, not a patch | Deterministic gate battery case study |
| Provide an interface that other engineering teams build evaluation logic on top of | Release-gate case study |
| Make an evaluation result reproducible so a re-run returns the same verdict | Artifact-lineage case study |
| Catch an incorrect pass or fail verdict from an automated evaluation before it ships | Deterministic gate battery case study |
What could go wrong hiring me
Most proof here is self-built, not run at a paid eval platform serving outside customers.
Counter: Release-gate case study
The evaluation corpus used is mine, not an independently curated benchmark.
Retrieval and indexing work here is at case-study scale, not managed-database scale.
Verify in 60 seconds
- Open the release-gate case study and read what one failing assertion does to a merge, /cases/eval-release-gate
- Open the deterministic gate battery case study and check its two rejected claims against the recomputed numbers, /cases/deterministic-gate-battery