Kubernetes Cost & Reliability Overhaul
$17,400/mo
saved
61%
node utilization
0
OOM incidents
A 40-node EKS cluster ran on static, hand-picked node groups with no autoscaling discipline, and pods were getting OOMKilled during every traffic spike. Compute spend kept climbing while reliability got worse, not better. Rightsized every workload's requests from real Prometheus usage data, then replaced static node groups and Cluster Autoscaler with Karpenter for just-in-time, bin-packed provisioning. Added PodDisruptionBudgets so consolidation never dropped a service below its minimum replica count.
Static m5.2xlarge node groups for every workload→Karpenter-provisioned nodes sized per pending pod
24% average CPU request utilization→61% average CPU request utilization
OOM kills on every traffic spike→Zero OOM incidents in 90 days post-launch