vlad.mironenko

Vlad Mironenko

Senior Infrastructure Engineer

I build the Kubernetes platform other engineers ship on.

Six years running large-scale EKS platforms — Terraform-managed infrastructure, GitOps delivery via Argo CD, and policy-as-code governance applied across every new application. Lately I build Claude Code skills that cut incident response time and operational toil.

Portrait of Vlad Mironenko

New York City

About

I operate Kubernetes platforms at scale and build the guardrails that let other teams ship on them safely.

At Epic Games I migrated 80+ applications off CI-driven Helm deploys onto a pull-based GitOps model on Argo CD, standardized policy-as-code across 12 EKS clusters with Kyverno and Policy Reporter, and brought event-driven autoscaling (KEDA) to the platform — cutting compute spend by roughly $15K/month by right-sizing capacity to actual demand instead of static headroom.

Before that, at Chowly, I rolled out Karpenter across three EKS clusters and pulled node provisioning time from ~5 minutes to under 1, extended Prometheus/Grafana with Thanos for global cross-cluster querying, and carried production on-call for the Kubernetes platform.

The last stretch of my work has been agentic: I build Claude Code skills that do real operational work — correlating incident telemetry with recent changes over read-only MCP, for example, which took RCA time on one recurring class of incident from ~50 minutes to ~10. The interesting problem isn't whether AI can touch infrastructure; it's building the guardrails so it does it safely.

Experience

Epic Games

Cary, NC
  • Senior Infrastructure Engineer
  • Migrated 80+ Kubernetes applications from CI-driven Helm deployments (GitHub Actions) to a pull-based GitOps model on Argo CD, reducing configuration drift across environments.
  • Implemented event-driven autoscaling with KEDA across the platform, reducing cluster compute spend by ~$15K/month by right-sizing capacity to real demand.
  • Standardized 12 EKS clusters with policy-as-code (Kyverno + Policy Reporter), establishing platform-wide deployment standards applied to all new applications while surfacing 200+ policy violations across existing workloads.
  • Consolidated legacy AWS infrastructure (IAM, RDS, S3) from CloudFormation to Terraform, importing 40+ previously unmanaged resources into state to bring the full estate under IaC.
  • Built a Claude Code skill for incident RCA that correlates Datadog telemetry with recent GitHub changes over read-only MCP, cutting investigation time from ~50 min to ~10 min.
  • Consolidated fragmented onboarding docs into a single guide, using it to onboard 4 engineers and reduce ramp-up time for future hires.

Chowly

Chicago, IL
  • Cloud Engineer
  • Rolled out Karpenter across 3 EKS clusters, dropping node provisioning time from ~5 min to under 1 min while scaling to 100+ nodes during peak traffic.
  • Extended Prometheus/Grafana with Thanos to provide global, deduplicated querying across clusters and long-term metric storage, giving teams a single source of truth for cross-cluster observability.
  • Automated cleanup of stale EBS snapshots and unused volumes with Python, eliminating idle storage costs and recurring manual maintenance.
  • Served in the 24/7 on-call rotation for production Kubernetes platforms, diagnosing and resolving infrastructure incidents across compute, networking, and deployments.

Selected work

Epic Games · platform-wide · 12 clusters

Policy-as-code across the EKS estate

  • Kyverno
  • Policy Reporter
  • EKS
  • Terraform

Standardized 12 EKS clusters on Kyverno-enforced policy, replacing tribal knowledge about what a "safe" deployment looks like with a checked rule set applied to every new application. Policy Reporter surfaced 200+ violations across existing workloads that had no other mechanism to be caught.

Epic Games · read-only MCP · ~50min → ~10min RCA

A Claude Code skill for incident root-cause analysis

  • Claude Code
  • MCP
  • Datadog
  • GitHub

Built a skill that correlates Datadog telemetry with recent GitHub changes over a read-only MCP connection — the same class of question an engineer asks first during triage, answered before a human opens a dashboard. Access is read-only by design: it accelerates investigation, and makes no change on its own.

Personal · own hardware · serves this page

A declarative on-prem private cloud

  • Harvester HCI
  • KubeVirt
  • RKE2
  • Terraform
  • FluxCD

On my own hardware, rebuilding a hand-managed homelab into a declarative on-prem private cloud: Harvester HCI for virtualization, RKE2 guest clusters provisioned from Terraform, reconciled by Flux from git. This site runs on that platform and is deployed the same way — a Kustomization Flux watches, an image pinned by digest, no port ever forwarded to serve it.

Skills

Cloud & Infrastructure

AWS (EC2, VPC, IAM, S3, RDS, Lambda, EKS) · Kubernetes · Karpenter · KEDA · Kyverno · Policy Reporter · Helm · Docker · Linux

IaC & Delivery

Terraform · GitHub Actions · Argo CD (GitOps) · CI/CD · Git

Observability

Datadog · Prometheus · Grafana · Thanos · CloudWatch · Alertmanager · PagerDuty

AI / Agentic

Claude Code · custom skills · MCP servers · context engineering

Security & Access

OIDC · External Secrets Operator / AWS Secrets Manager

Scripting

Bash · Python

Education

Master of Science, System Analysis and Management SSU, Krasnoyarsk, Russia
Bachelor of Science, Information and Communication Technology SSU, Krasnoyarsk, Russia
  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • HashiCorp Certified: Terraform Associate (003)

Contact

Open to conversations about senior platform and infrastructure roles — especially teams treating AI-assisted operations as an engineering problem, not a demo.