Skip to main content

Michael Puodziukas

AI Evaluation Engineer

Remote, US-based (Mountain Time) - michael@puodziukas.dev - puodziukas.dev - github.com/mpuodziukas-labs

Role fit

Mid-level to senior, agentic AI + AI security

Availability

Open to mid-level and senior IC roles - W2, remote (US).

Outcome

Public probe reproduces a failure class standard audits miss - clone-and-run proof

Summary

Independent engineer, building open-source AI evaluation and security tooling since 2023. I build the gates that catch a bad AI output before it ships, and the artifacts that let a stranger re-run the result: the adversarial-validation gate below catches 140/140 attack attempts on its own synthetic test corpus (n=220, written for the repo), 0 false positives. Every result ships as clone-and-run code.

One operator, three disciplines usually spread across three hires: agentic-AI engineering, OWASP-mapped adversarial security, and COBOL arithmetic forensics (the probe below). Each ships with a public, clone-and-run replay; the combination is the differentiator, not any single line.

Experience

Independent AI Research & Engineering

Remote - 2023 - Present

Build AI systems and red-team them as the same person; every result ships as clone-and-run code.

Agentic AI systems: design, build, and adversarial validation

  • Stood up an evaluation harness with a golden replay corpus so routing and generation decisions are re-executable and auditable; regressions are caught before release, not after.
  • Implemented MLX kernel optimization for the home-lab self-hosted inference stack: my public fork of exo (open-source distributed inference), with an ops-hardening branch (mpuodziukas-labs/exo).

AI security: LLM red-teaming and guardrails

  • Authored an open-source adversarial gate (llm-adversarial-gate): OWASP LLM Top 10 coverage; the gate is a deterministic, rule-based check, not a model call. Zero misses on its own synthetic test corpus (n=220, written for the repo): 140 adversarial + 80 benign, 140/140 caught, 0/80 over-blocked. Clone and run.
  • Designed defense-in-depth for indirect prompt injection and agent hijacking; each defense paired with an adversarial replay case so coverage is demonstrable, not asserted.

Mainframe forensics: COBOL arithmetic integrity

  • Published an open-source COBOL arithmetic forensic tool (cobol-pic-probe): detects COMP-3 / PIC-clause precision defects that silently corrupt financial arithmetic. Clone and run.

Infrastructure: self-hosted inference (home lab)

  • Built and operate a self-hosted inference cluster at home; the build is public: my public fork of exo (open-source distributed inference), with an ops-hardening branch (github.com/mpuodziukas-labs/exo).
  • Model placement is pipeline-sharded across nodes with a live admission check on real available memory, not static config, so the cluster serves the full model rather than a small fallback.
  • Instrumented everything with deterministic replay and CI-gated assertions so failures surface as alerts, not incidents.

Verifiable Proof (clone & run)

Technical Skills

Languages: Python - Rust - TypeScript - C - SQL - Bash - COBOL

Agentic / LLM: Multi-agent orchestration - LangChain - deterministic state engines - evaluation harnesses - RAG-freshness engineering (embedding-drift oracle - retrieval-recall floor)

AI Ops / Reliability: drift & health oracles for the home-lab inference cluster - ratchet loops (research, critique, rewrite, score, repeat until a deterministic oracle passes) - mutation-proven gates (a planted failure must flip green to red)

AI Security: LLM red-teaming - OWASP LLM Top 10 - prompt-injection & agent-hijacking defense - adversarial replay corpora

Inference / ML: MLX - vLLM - ONNX - quantization (4-bit and 8-bit) - tensor & pipeline parallelism (Apple Silicon)

Data / Streaming: Kafka - Redis - Qdrant - deterministic replay

Infra / SRE: Docker - Terraform - OpenTelemetry - CI/CD

Mainframe / Legacy: COBOL - COMP-3 forensics - COBOL to Rust byte-exact migration

Security / Cryptography: GPG-signed checksum manifests (home lab, byte-identical verification at scale)

AI Governance (engineering artifacts, not compliance opinion): EU AI Act high-risk system requirements - NIST AI RMF-aligned documentation

Education

University of South Florida, B.S. Entrepreneurship

Western Governors University, Computer Science coursework (no degree)

Relevant coursework (incl. transfer credit): Calculus I-II, Discrete Math I, Trigonometry, Statistics I-II, Data Structures and Algorithms I, Version Control, intro Python and Java, Web Development, Networking, Relational Databases, Database Programming and Management. Self-study: Discrete Math II, Linear Algebra, DSA II, software QA, Harvard CS50 (no certificate), Helsinki Java MOOC.

Availability

Open to mid-level and senior IC roles - W2, remote (US). Domains: agentic AI, AI security / red-teaming, inference infrastructure, mainframe modernization.