The Cast AI blog
Guides, tutorials, and tips on Kubernetes automation, from cost optimization to cloud security and everything in between.
Cast AI vs SkyPilot: Two Approaches to Finding GPUs Across Clouds
SkyPilot and Cast AI take different approaches to GPU infrastructure, and the right fit depends on the workload. This post compares SkyPilotās multi-cloud job orchestration and…

Kubernetes Configuration Drift: How to Detect It and What It Costs You
Kubernetes configuration drift occurs when actual cluster state diverges from what Git or Helm declares,…

GPU Job Queueing with Kueue and DRA: Scheduling AI Workloads Without Idle Capacity
Kueue and DRA help Kubernetes manage GPU capacity more efficiently by queuing workloads until resources…

Migrating from Kubecost to Automated Optimization: What Changes and What to Keep
Moving from Kubecost to Cast AI is mainly about planning the transition. Keep your existing…

How Cast AI’s Automation Decides: The Guardrails, Rollbacks and Evidence Behind Each Action
Safe Kubernetes automation requires more than utilization metrics. This post explores how SLO signals, pre-flight…

Reducing Kubernetes Toil: A Ranked List of What to Automate First
Kubernetes toil grows with cluster scale, but not every task is equally worth automating. This…

What Is AI SRE, and Where Does Cost Automation Fit?
AI SRE goes beyond observability and AIOps into incident investigation, root cause analysis, and remediation.…

Automated, Autonomous, Agentic: What the Three Levels of Kubernetes Operations Actually Mean
Automation, autonomy, and agentic operations are not interchangeable. This post breaks down the four levels…

The FOCUS Specification for Kubernetes and AI Costs, Explained
FOCUS 1.4 standardizes cloud billing data across providers and improves Kubernetes cost visibility, but does…

Kubernetes Chargeback and Showback: How to Bill Teams for What They Actually Use
Kubernetes chargeback only works when teams trust the numbers behind the bill. This guide explains…
