I operate Kubernetes platforms at scale and build the guardrails that let other teams ship on them safely.
At Epic Games I migrated 80+ applications off CI-driven Helm deploys onto a
pull-based GitOps model on Argo CD, standardized policy-as-code across 12 EKS
clusters with Kyverno and Policy Reporter, and brought event-driven autoscaling
(KEDA) to the platform — cutting compute spend by roughly $15K/month by
right-sizing capacity to actual demand instead of static headroom.
Before that, at Chowly, I rolled out Karpenter across three EKS clusters and
pulled node provisioning time from ~5 minutes to under 1, extended
Prometheus/Grafana with Thanos for global cross-cluster querying, and carried
production on-call for the Kubernetes platform.
The last stretch of my work has been agentic: I build Claude Code skills that
do real operational work — correlating incident telemetry with recent changes
over read-only MCP, for example, which took RCA time on one recurring class of
incident from ~50 minutes to ~10. The interesting problem isn't whether AI can
touch infrastructure; it's building the guardrails so it does it safely.