The conservative engineering culture that makes Azure Kubernetes Service (AKS) the default choice for enterprise infrastructure is the exact reason those same clusters now sit on the worst utilization metrics in cloud compute.
A separate post looked at the problem – AKS clusters are carrying substantial unused capacity across CPU, memory, and GPUs. The obvious next question is what to actually change.
There isn’t one setting that fixes this. Cost optimization happens at several layers, and those decisions affect each other. Workload requests influence node provisioning. Node selection affects consolidation and Spot opportunities. GPU allocation determines how much physical hardware is required. And the resulting infrastructure footprint determines how much committed capacity makes sense.
Here is what those decisions look like in practice.
Node Provisioning: NAP vs. Cluster Autoscaler
Cluster Autoscaler scales existing node pools, with each pool tied to a particular VM SKU. That works well for predictable workloads, but it can leave capacity stranded.
If a node pool contains Standard_D4s_v3 instances with 4 vCPUs and 16 GiB of memory, a pod requesting 20 GiB cannot use it. You either provision a larger pool or leave the pod pending.
Node Auto Provisioning (NAP), which reached general availability on AKS in July 2025, takes a different approach. Built on Karpenter, it evaluates pending pod requirements and selects an appropriate VM from the instance types allowed by the NodePool.
A minimal configuration looks like this:
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
name: general-purpose
spec:
template:
spec:
nodeClassRef:
group: karpenter.azure.com
kind: AKSNodeClass
name: default
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
- key: karpenter.azure.com/sku-family
operator: In
values: ["D", "E"]
disruption:
consolidationPolicy: WhenEmptyOrUnderutilized
consolidateAfter: 120s
limits:
cpu: 1000
---
apiVersion: karpenter.azure.com/v1beta1
kind: AKSNodeClass
metadata:
name: default
spec:
imageFamily: AzureLinux
osDiskSizeGB: 128Check the AKSNodeClass API version when migrating older configurations. karpenter.azure.com/v1beta1 is the current version; v1alpha2 is deprecated.
The other setting worth paying attention to is consolidateAfter. A value of 120 seconds is a reasonable starting point if you want to give short-lived traffic drops time to recover before consolidation. Tune it against the workload’s actual behavior rather than treating two minutes as a universal production setting.
Because NAP can consolidate more aggressively than Cluster Autoscaler, PodDisruptionBudgets also matter. Audit them before migrating production workloads.
CA remains useful for workloads with specific VM, hardware, or licensing requirements, or where more conservative disruption behavior is important. CA and NAP can also coexist, making a gradual migration possible.
In practice: Caudalie
If we look at some of the examples in the 2026 State of Kubernetes Optimization Report, Caudalie is a good illustration of what this looks like in practice. The team had already implemented extensive cost and performance optimizations, but still needed a better way to handle traffic fluctuations without carrying excess capacity.
Caudalie introduced automated autoscaling that provisions capacity in line with demand. In the period shown in the report, the cluster’s provisioned capacity drops sharply as the autoscaler takes over, while the workload continues to run.
The point isn’t simply that the cluster got smaller. It’s that the infrastructure was able to respond to actual demand instead of remaining sized around a previous assumption.
Pod Rightsizing with VPA
NAP can provision the right node for a workload, but it cannot tell whether the workload’s resource requests are themselves oversized.
Vertical Pod Autoscaler (VPA) addresses that problem by comparing declared requests with observed consumption. Its operating modes range from recommendation-only Off mode through Initial and Recreate, to InPlaceOrRecreate, which attempts to change resources without restarting the pod.
InPlaceOrRecreate is available on AKS 1.34+. It allows VPA to adjust pod resources in place where supported, falling back to recreation when necessary. Verify support for the specific AKS version before using it in production.
For in-place CPU resizing, the container needs an explicit resize policy:
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
spec:
template:
spec:
containers:
- name: app
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "2000m"
memory: "2Gi"
resizePolicy:
- resourceName: cpu
restartPolicy: NotRequired
- resourceName: memory
restartPolicy: RestartContainerNotRequired allows CPU resources to be resized without restarting the container. Memory resizing generally still requires a restart.
There is also an important HPA interaction. Avoid running automated VPA modes alongside an HPA that uses CPU or memory as its scaling signal. VPA changes the resource request, which changes the basis of the HPA calculation and can create an unwanted feedback loop. If both are required, use application-level HPA signals such as requests per second, queue depth, or latency.
In practice: Mercedes-Benz.io
Mercedes-Benz.io found that resource requests across its large Kubernetes footprint no longer reflected actual consumption. The team used automation to adjust requests based on observed usage rather than manually reviewing deployments.
The benefit wasn’t necessarily one dramatically oversized workload. It was the accumulation of small mismatches across a large fleet.
The report also shows why rightsizing isn’t necessarily a reliability trade-off. In one production cluster, OOM kills were occurring 40 to 50 times per measurement interval despite generous resource padding. After automated rightsizing was enabled, OOM kills dropped to near zero while provisioned CPU was cut roughly in half.
Giving every workload more resources wasn’t solving the reliability problem. Identifying which workloads actually needed more resources did.
Spot VMs: Cheap Capacity, Managed Properly
Azure Spot VMs can provide discounts of up to 90% compared with pay-as-you-go pricing, but the capacity can be reclaimed with as little as 30 seconds’ notice.
That makes application configuration important:
apiVersion: apps/v1
kind: Deployment
metadata:
name: batch-processor
spec:
template:
spec:
terminationGracePeriodSeconds: 25
tolerations:
- key: "kubernetes.azure.com/scalesetpriority"
operator: "Equal"
value: "spot"
effect: "NoSchedule"
containers:
- name: processor
image: my-batch-app:latestSetting terminationGracePeriodSeconds to 25 seconds or less leaves some margin inside the 30-second eviction window. Pair Spot workloads with a PodDisruptionBudget where multiple replicas need to remain available.
Spot is generally appropriate for batch processing, CI runners, suitable stateless services, and ML workloads that can checkpoint and recover. It is not appropriate for stateful services without fault tolerance or applications that cannot tolerate node replacement.
You also need a fallback. With Cluster Autoscaler, preferred and fallback pools can be configured through the priority expander. With NAP, NodePools can allow both Spot and on-demand capacity so workloads can fall back when suitable Spot capacity isn’t available.
GPUs: Share the Hardware Before Buying More
GPU workloads make the same problem considerably more expensive. Many inference workloads don’t continuously saturate their GPUs, but dedicated deployment models still reserve an entire GPU for each workload.
Time-slicing and GPU sharing allow multiple workloads to use the same physical hardware. On supported NVIDIA GPUs, Multi-Instance GPU (MIG) can go further by partitioning a physical GPU into isolated GPU instances. The right approach depends on the workload and the level of isolation it requires.
The important question is whether a workload actually needs a dedicated physical GPU or simply needs access to GPU capacity. If the latter is true, sharing and better scheduling can reduce the amount of physical GPU infrastructure required without reducing the capacity available to workloads when they need it.
Commitments: Optimize First, Commit Second
Reserved Instances and Savings Plans can significantly reduce the cost of predictable baseline capacity. But the baseline needs to be right first.
A one-year Azure Reserved VM Instance saves approximately 48% compared with pay-as-you-go pricing, while three-year reservations can reach up to 72%. Azure Savings Plans provide more flexibility, with savings of up to 65% at the three-year tier.
The mistake is committing before rightsizing.
If workloads are running on Standard_D8s_v3 because their resource requirements were set conservatively, then later turn out to fit comfortably on D4 instances, a three-year commitment to the larger SKU has simply turned overprovisioning into a long-term financial commitment.
The sequence should be:
- Measure actual workload consumption.
- Rightsize pod requests.
- Optimize node provisioning and consolidation.
- Move appropriate workloads to Spot.
- Commit to the remaining baseline.
Following that sequence means the commitment is based on the infrastructure you actually need, rather than today’s inflated resource requests and historical node configuration.
Don’t Ignore ARM
The underlying VM architecture is another lever worth evaluating. ARM CPU nodes grew 3.5 times faster than x86 between Q2 2024 and Q4 2025, and ARM now accounts for about 9% of the CPU fleet analyzed in the report.
On AKS, the Dpls_v5 series, powered by Ampere Altra processors, is another option worth testing for suitable cloud-native workloads. The actual price-performance benefit depends on the workload, so benchmark before expanding beyond a small set of services.
Before moving workloads, verify that container images support linux/arm64, check for compiled binaries that assume x86, and identify node-level dependencies tied to the architecture.
There’s no need to migrate an entire cluster. Add an ARM node pool, route a few suitable stateless workloads to it, and compare performance and cost before expanding.
Fixing AKS Utilization Is Possible When Done Right
The examples in the report are useful because they show that this isn’t just a problem you see in small or poorly managed clusters. The Mercedes-Benz.io rightsizing example and the Caudalie autoscaling example both show substantial room to improve how infrastructure is being used, even in production environments.
The problems are familiar and surprisingly common anti-patterns like workloads asking for more than they use, nodes carrying more capacity than they need, GPUs sitting idle between bursts, and commitments that don’t always match actual demand. At massive fleet scale, those small mismatches become very expensive.
But don’t panic. The examples also show that the utilization gap is not inevitable and can be reduced. Even across large and complex Kubernetes environments, these changes can be made safely when they are approached systematically: understand what is actually being used, make one change at a time, measure its impact, and use that information to inform the next decision.
For AKS teams, that means there is no reason to look at that 2% GPU utilization or the 69% CPU overprovisioning and conclude that fixing it requires a wholesale infrastructure overhaul. There are multiple places to start, and each improvement gives you a better baseline for the next one.
The key is not to optimize everything at once. The DevOps mindset applies here too: start with a small, controlled change, make sure it behaves as expected in production, and then roll it out more broadly. That’s how you improve utilization without turning cost optimization into a risky fleet-wide exercise.
Cut Kubernetes costs with automation
- Rightsize workloads
- Reduce overprovisioning
- Scale clusters efficiently



