Michael Puodziukas
AI Evaluation Engineer
Mid-level to senior, agentic AI + AI security
Open to mid-level and senior IC roles - W2, remote (US).
Public probe reproduces a failure class standard audits miss - clone-and-run proof
Summary
Independent engineer, building open-source AI evaluation and security tooling since 2023. I build the gates that catch a bad AI output before it ships, and the artifacts that let a stranger re-run the result: the adversarial-validation gate below catches 140/140 attack attempts on its own synthetic test corpus (n=220, written for the repo), 0 false positives. Every result ships as clone-and-run code.
One operator, three disciplines usually spread across three hires: agentic-AI engineering, OWASP-mapped adversarial security, and COBOL arithmetic forensics (the probe below). Each ships with a public, clone-and-run replay; the combination is the differentiator, not any single line.
Experience
Independent AI Research & Engineering
Build AI systems and red-team them as the same person; every result ships as clone-and-run code.
Agentic AI systems: design, build, and adversarial validation
- Stood up an evaluation harness with a golden replay corpus so routing and generation decisions are re-executable and auditable; regressions are caught before release, not after.
- Implemented MLX kernel optimization for the home-lab self-hosted inference stack: my public fork of exo (open-source distributed inference), with an ops-hardening branch (mpuodziukas-labs/exo).
AI security: LLM red-teaming and guardrails
- Authored an open-source adversarial gate (llm-adversarial-gate): OWASP LLM Top 10 coverage; the gate is a deterministic, rule-based check, not a model call. Zero misses on its own synthetic test corpus (n=220, written for the repo): 140 adversarial + 80 benign, 140/140 caught, 0/80 over-blocked. Clone and run.
- Designed defense-in-depth for indirect prompt injection and agent hijacking; each defense paired with an adversarial replay case so coverage is demonstrable, not asserted.
Mainframe forensics: COBOL arithmetic integrity
- Published an open-source COBOL arithmetic forensic tool (cobol-pic-probe): detects COMP-3 / PIC-clause precision defects that silently corrupt financial arithmetic. Clone and run.
Infrastructure: self-hosted inference (home lab)
- Built and operate a self-hosted inference cluster at home; the build is public: my public fork of exo (open-source distributed inference), with an ops-hardening branch (github.com/mpuodziukas-labs/exo).
- Model placement is pipeline-sharded across nodes with a live admission check on real available memory, not static config, so the cluster serves the full model rather than a small fallback.
- Instrumented everything with deterministic replay and CI-gated assertions so failures surface as alerts, not incidents.
Verifiable Proof (clone & run)
- github.com/mpuodziukas-labs/cobol-pic-probe: COBOL COMP-3 / PIC precision-defect forensics.
- github.com/mpuodziukas-labs/llm-adversarial-gate: OWASP LLM Top 10-mapped, rule-based guardrail gate.
Technical Skills
Languages: Python - Rust - TypeScript - C - SQL - Bash - COBOL
Agentic / LLM: Multi-agent orchestration - LangChain - deterministic state engines - evaluation harnesses - RAG-freshness engineering (embedding-drift oracle - retrieval-recall floor)
AI Ops / Reliability: drift & health oracles for the home-lab inference cluster - ratchet loops (research, critique, rewrite, score, repeat until a deterministic oracle passes) - mutation-proven gates (a planted failure must flip green to red)
AI Security: LLM red-teaming - OWASP LLM Top 10 - prompt-injection & agent-hijacking defense - adversarial replay corpora
Inference / ML: MLX - vLLM - ONNX - quantization (4-bit and 8-bit) - tensor & pipeline parallelism (Apple Silicon)
Data / Streaming: Kafka - Redis - Qdrant - deterministic replay
Infra / SRE: Docker - Terraform - OpenTelemetry - CI/CD
Mainframe / Legacy: COBOL - COMP-3 forensics - COBOL to Rust byte-exact migration
Security / Cryptography: GPG-signed checksum manifests (home lab, byte-identical verification at scale)
AI Governance (engineering artifacts, not compliance opinion): EU AI Act high-risk system requirements - NIST AI RMF-aligned documentation
Education
University of South Florida, B.S. Entrepreneurship
Western Governors University, Computer Science coursework (no degree)
Relevant coursework (incl. transfer credit): Calculus I-II, Discrete Math I, Trigonometry, Statistics I-II, Data Structures and Algorithms I, Version Control, intro Python and Java, Web Development, Networking, Relational Databases, Database Programming and Management. Self-study: Discrete Math II, Linear Algebra, DSA II, software QA, Harvard CS50 (no certificate), Helsinki Java MOOC.
Availability
Open to mid-level and senior IC roles - W2, remote (US). Domains: agentic AI, AI security / red-teaming, inference infrastructure, mainframe modernization.