Back to Insights
Cloud

Kubernetes Cost Optimization: A Practical Guide to Cutting Cloud Spend

11 min read

Most Kubernetes clusters run at under 25% utilisation. Here is the ordered sequence of changes that reclaims spend without introducing incidents.

Kubernetes does not make infrastructure expensive; it makes waste invisible. Because the scheduler bills you for reserved capacity rather than consumed capacity, a cluster can sit at 20 percent real utilisation while every node reports as full. The fix is not a tool purchase — it is a sequence of changes applied in the right order, from lowest risk to highest.

Understand where the money actually goes

Before changing anything, decompose spend into four buckets: compute (nodes), storage (volumes and snapshots), network (egress and cross-zone traffic), and managed services (load balancers, registries, control planes). In typical clusters compute is 60 to 75 percent of the bill, but cross-zone traffic and orphaned volumes are the two line items teams consistently underestimate.

You cannot optimise what you cannot attribute. Enforce namespace and label conventions for team, service, and environment, then produce a weekly cost-per-service report. Attribution changes behaviour faster than any policy document.

Step one: right-size requests and limits

This is the highest return, lowest risk change available. Pull 14 days of CPU and memory usage per container and set requests near the 95th percentile of observed usage with a modest buffer, not at the peak of the worst day.

  • CPU: set requests from the p95, and leave limits generous or unset for latency-sensitive services — CPU limits cause throttling that looks like application latency.
  • Memory: set requests and limits close together, since memory is not compressible and overcommitment ends in OOM kills.
  • Automate the loop. A vertical autoscaler in recommendation mode gives you continuous right-sizing data without surprise restarts.

Expect to reclaim 30 to 50 percent of node capacity in a cluster that has never done this.

Step two: make scaling actually elastic

Right-sizing frees capacity; autoscaling turns freed capacity into money saved. Three layers need to work together:

LayerMechanismCommon mistake
Pod countHorizontal pod autoscalerScaling on CPU when the real constraint is queue depth or concurrency
Pod sizeVertical pod autoscalerRunning it in auto mode alongside an HPA on the same metric
Node countCluster autoscaler or Karpenter-style provisionerBlocking scale-down with pods that have no disruption budget

Scale-down is where most clusters fail. Audit for pods that pin nodes open: unreplicated deployments, missing pod disruption budgets, local storage dependencies, and long terminationGracePeriodSeconds values.

Step three: buy capacity correctly

Match commitment to workload stability. Baseline capacity that runs continuously belongs on committed-use or reserved pricing — typically 30 to 55 percent cheaper. Bursty, fault-tolerant workloads belong on spot capacity at 60 to 90 percent discounts. The discipline that makes spot safe is diversification across several instance families and availability zones, plus handlers that drain gracefully on the interruption notice.

Keep on-demand for stateful databases, single-replica services, and anything whose restart cost exceeds the savings.

Step four: fix storage and network leakage

  • Orphaned volumes. Persistent volume claims outlive the workloads that created them; sweep monthly.
  • Over-provisioned IOPS. Most services do not need premium tiers; verify against actual latency requirements.
  • Cross-zone chatter. Enable topology-aware routing so traffic prefers same-zone endpoints. This is frequently a five-figure annual line item that nobody owns.
  • Log volume. Sample high-cardinality debug logging in production; observability bills often rival compute.

Step five: make efficiency a standing practice

Optimisation projects regress within two quarters unless they become operational habit. Three mechanisms hold the gains: a cost-per-service dashboard reviewed in the same meeting as reliability metrics; admission policy that rejects workloads without requests, limits, and ownership labels; and a monthly review of the ten most expensive services with a named owner per line.

Set the target as efficiency, not absolute spend. Cost per transaction or per active user keeps the goal aligned with growth rather than punishing it.

What not to do

Avoid aggressive bin-packing that removes headroom for failure domains, cutting replica counts below what your availability target requires, and chasing marginal savings on services that represent two percent of the bill. The reliability cost of over-optimisation always shows up eventually, and it is more expensive than the compute it saved.

Frequently asked questions

Why are Kubernetes clusters so expensive?

Because you pay for requested capacity, not used capacity. Generous CPU and memory requests get reserved by the scheduler, so a cluster can be fully allocated while actual utilisation sits near 20 percent.

What is the fastest way to reduce costs?

Right-size requests from observed usage percentiles. It typically reclaims 30 to 50 percent of node capacity and requires no architectural change.

Are spot instances safe in production?

Yes for stateless, replicated, interruption-tolerant workloads with disruption budgets and diversified instance types. Keep stateful and single-replica services on on-demand or reserved capacity.

How much can a typical team save?

Clusters that have never been optimised commonly reduce compute spend 35 to 60 percent within a quarter by combining right-sizing, working scale-down, and correct commitment purchasing.

Tagged With:

Kubernetes
FinOps
cloud cost
autoscaling
DevOps

Ready to Transform Your Digital Experience?

Let's discuss how Kinematic Digital can help you achieve your business goals.