Engineering
Explore technical articles from Cast AI covering Kubernetes engineering, cloud infrastructure, automation and platform development.

GPU sharing and GPU Rightsizing: Stop Wasting GPU
Learn how to reduce Kubernetes GPU waste when average utilization sits at just 5%. This…

GPU Scheduling and Bin-Packing in Kubernetes: Pack More AI onto Every GPU
Learn how Kubernetes GPU scheduling affects utilization, cost, and AI workload density. This guide covers…

Kubernetes Bin-Packing and Node Consolidation: How to Cut Idle Node Cost
69% of Kubernetes clusters over-provision CPU. The default scheduler is partly to blame. Here is…

Kubernetes GPU Autoscaling: Scale GPU Capacity to Real Demand
GPU autoscaling in Kubernetes matches GPU capacity to real AI workload demand. Learn to configure…

KEDA: How Event-Driven Autoscaling Cuts Kubernetes Cost
Learn how KEDA enables event-driven autoscaling in Kubernetes using external metrics such as Kafka, Prometheus,…

HPA vs VPA: When to Use Each, and Can You Use Both?
HPA scales the number of pod replicas; VPA scales the CPU and memory each pod…

A CTO’s Guide to Kubernetes Cost Optimization
Learn how the Kubernetes Optimization Loop helps engineering teams reduce infrastructure costs while improving unit…

Kubernetes Cost Optimization Checklist: 37 Checks Before You Touch Production
Learn how to reduce Kubernetes costs with a systematic optimization strategy that addresses overprovisioning, idle…

Cast AI vs Kubecost: Cost Visibility or Automated Optimization?
Compare Kubecost and CAST AI to understand the difference between Kubernetes cost visibility and automated…