Michael Puodziukas
Senior / Staff / Principal Engineer · Agentic AI · AI Red-Teaming · Inference Infrastructure
Staff / Principal, agentic AI + AI security
Builds and red-teams the same systems, 3+ yrs
Q3 2026 · remote · no relocation
Inference bill cut to $0/yr API · found a defect 3 years of audits missed
Summary
Staff/Principal engineer who builds agentic AI systems and red-teams them as the same person, so the failure class surfaces in the build, not in production. Eliminated a $26,400/mo inference bill on a self-hosted 158 TPS cluster; separately migrated a COBOL batch core to Rust for 27× throughput; drove a measured hallucination rate from 27% to under 1%. Every result ships as clone-and-run code. Remote, available Q3 2026.
One operator, three disciplines that are usually three hires: agentic-AI engineering, OWASP/MITRE ATLAS-mapped adversarial security, and financial-system forensics (the COBOL defect above). Each ships with a public, clone-and-run replay; the combination is the differentiator, not any single line.
Experience
Independent AI Research & Engineering
Build production AI systems and red-team them as the same person; every result ships as clone-and-run code.
Agentic AI systems — design, build, and adversarial validation
- Built the multi-agent orchestration behind the 27%-to-under-1% hallucination cut: deterministic state, confidence gating, and a 3-lens adversarial council (correctness · citation · reproducibility) that blocks any single-agent output from shipping unchecked.
- Stood up an evaluation harness with a golden replay corpus so routing and generation decisions are re-executable and auditable; regressions are caught before release, not in production.
- Cut agent cold-start to ~200ms p99 through MLX kernel optimization and speculative decoding on Apple Silicon.
AI security — LLM red-teaming and guardrails
- Authored an open-source adversarial gate (llm-adversarial-gate): OWASP LLM Top 10 and MITRE ATLAS-mapped coverage; the gate is a deterministic check, not a model call. Zero misses across the public adversarial corpus: 140 adversarial + 80 benign = 220, 140/140 caught, 0/80 over-blocked. Clone and run.
- Designed defense-in-depth for indirect prompt injection, agent hijacking, and MCP tool-schema validation; each defense paired with an adversarial replay case so coverage is demonstrable, not asserted.
Mainframe modernization — COBOL → Rust, byte-exact
- Published an open-source COBOL arithmetic forensic tool (cobol-pic-probe): detects COMP-3 / PIC-clause precision defects that silently corrupt financial arithmetic. Clone and run.
- Migrated a COBOL batch core to idiomatic Rust with byte-exact arithmetic verified by deterministic replay. The throughput gain above held with zero behavioral regressions, validated across 10M fuzz iterations with zero panics.
Infrastructure / SRE — sovereign inference and observability
- Built and operate the 2-node sovereign Apple Silicon cluster behind that inference bill (M1 Max 32GB + M4 Mini 16GB = 48GB unified) over Thunderbolt 4 zero-copy transport, holding 389ms P50 end-to-end entirely on-prem with zero cloud API dependency; the build is public: 55 commits ahead of upstream on a production fork of exo (github.com/mpuodziukas-labs/exo).
- Took a $26,400/mo managed-API bill to $0/yr API via a model-right-sizing + KV-cache + prefix-caching stack, holding SLAs. Hardware and power still apply; no total-cost-zero claim.
- Ring placement is pipeline-sharded across both nodes with a live admission check on real available memory, not static config: the cluster serves a 30B Q4-class model split across the full ring, not a single node running a fallback 3B.
- Instrumented everything with Prometheus-style telemetry, deterministic replay, and CI-gated assertions so failures surface as alerts, not incidents.
Verifiable Proof (clone & run)
- github.com/mpuodziukas-labs/cobol-pic-probe: COBOL COMP-3 / PIC precision-defect forensics.
- github.com/mpuodziukas-labs/llm-adversarial-gate: OWASP LLM Top 10 / MITRE ATLAS-mapped guardrail gate.
Technical Skills
Languages: Python · Rust · Go · TypeScript · Java · C++ · C · Scala · SQL · Bash · COBOL II
Agentic / LLM: Multi-agent orchestration · LangGraph · LangChain · deterministic state engines · evaluation harnesses · RAG-freshness engineering (embedding-drift oracle · retrieval-recall floor · idempotent re-embed) · confidence gating · KV-cache & prefix caching · speculative decoding
AI Ops / Reliability: self-healing inference substrate (drift & health oracles, auto-recovery daemons) · ratchet loops (research → critique → rewrite → score, repeat until a deterministic oracle passes) · mutation-proven gates (a planted failure must flip green → red) · sensitivity-first model routing (sensitive → local fail-closed, premium reserved for judgment)
AI Security: LLM red-teaming · OWASP LLM Top 10 · MITRE ATLAS · prompt-injection & agent-hijacking defense · MCP tool-schema validation · adversarial replay corpora
Inference / ML: MLX · PyTorch · vLLM · Ray · Triton · ONNX · CUDA · quantization (AWQ/GPTQ, 4/8-bit) · speculative decoding · tensor & pipeline parallelism · multi-architecture serving (Apple Silicon · x86 · CUDA)
Data / Streaming: Kafka · Postgres · Redis · Qdrant · event sourcing · deterministic replay
Infra / SRE: Docker · Terraform · Prometheus · OpenTelemetry · zero-trust mTLS · CI/CD
Cloud → Sovereign: model-agnostic inference layers · cloud-to-sovereign migration (zero-egress, air-gappable)
Mainframe / Legacy: COBOL II · COMP-3 forensics · COBOL→Rust byte-exact migration
Security / Cryptography: FIPS 203/204/205 (PQC) · SHA-256 audit trails
AI Governance (engineering artifacts, not compliance opinion): EU AI Act high-risk system requirements · NIST AI RMF-aligned documentation · SHA-256 hash-chained audit trails for LLM decisions · kill-switch architecture
Education
University of South Florida — B.S. Business Administration & Entrepreneurship, Muma College of Business
Western Governors University — B.S. Computer Science (in progress)
Availability
Remote · Senior / Staff / Principal IC · available Q3 2026 · compensation negotiated by role and scope · domains: agentic AI, AI security / red-teaming, inference infrastructure, mainframe modernization.