Skip to main content

puodziukas.dev

AI Evaluation Engineer

Open to mid-level and senior IC roles, W2, remote (US).

I build the systems that run evaluation tests and check that their results are correct.

eval runcontrol PASSplanted regression FAILgatereceipt

Requirements and proof

RequirementProof
Build the backend platform that executes evaluation tests and processes their resultsRelease-gate case study
Improve the correctness, performance, and observability of a test-execution systemRelease-gate case study - Deterministic gate battery case study
Investigate a live issue to root cause and ship a durable fix, not a patchDeterministic gate battery case study
Provide an interface that other engineering teams build evaluation logic on top ofRelease-gate case study
Make an evaluation result reproducible so a re-run returns the same verdictArtifact-lineage case study
Catch an incorrect pass or fail verdict from an automated evaluation before it shipsDeterministic gate battery case study

What could go wrong hiring me

Most proof here is self-built, not run at a paid eval platform serving outside customers.

Counter: Release-gate case study

The evaluation corpus used is mine, not an independently curated benchmark.

Counter: Deterministic gate battery case study

Retrieval and indexing work here is at case-study scale, not managed-database scale.

Counter: Artifact-lineage reproducibility case study


Verify in 60 seconds

  1. Open the release-gate case study and read what one failing assertion does to a merge, /cases/eval-release-gate
  2. Open the deterministic gate battery case study and check its two rejected claims against the recomputed numbers, /cases/deterministic-gate-battery

Work with meFull hire pageRésumé (PDF)

puodziukas.dev