For recruiters and hiring managers hiring a AI Evaluation Analyst (evidence review, test reporting).
Role fit, with proof you can run
AI Evaluation Analyst (evidence review, test reporting)
Mid-level and senior IC roles. W2 full-time or contract-to-hire, remote US (Arizona).
I read evaluation results, check each verdict against its evidence, and report what passed, what failed and why. The same public proofs back the engineer lane.

eval resultscheck evidenceverdictreportreceipt
Requirements and proof
| Requirement | Proof |
|---|---|
| Read automated evaluation results and report which checks passed, failed or were skipped | Release-gate case study |
| Check a pass or fail verdict against its evidence before it is reported | Deterministic gate battery case study |
| Trace a reported result back to the exact artifact version that produced it | Artifact-lineage case study |
| Re-run a test and confirm the same verdict comes back | Public RAG groundedness repo |
| Write up a defect and its root cause so another team can act on it | Deterministic gate battery case study |
What could go wrong hiring me
Most proof here is self-built, not reporting for a team that depends on the results.
Counter: Release-gate case study
The evaluation corpus used is mine, not an independently curated benchmark.
My reporting samples are case studies, not dashboards used by a business team.
Counter: Artifact-lineage case study
Verify in 60 seconds
- Open the release-gate case study and read what one failing assertion does to a merge, /cases/eval-release-gate
- Open the deterministic gate battery case study and check its two rejected claims against the recomputed numbers, /cases/deterministic-gate-battery