,

Cast AI vs SkyPilot: Two Approaches to Finding GPUs Across Clouds

SkyPilot and Cast AI take different approaches to GPU infrastructure, and the right fit depends on the workload. This post compares SkyPilot’s multi-cloud job orchestration and neocloud access with Cast AI’s managed Kubernetes optimization, GPU sharing, and cost management for long-running inference services.

cast ai vs skypilot featured image

SkyPilot and Cast AI both solve the problem of GPUs being unavailable where you want them, from different starting points. SkyPilot is an open-source framework that launches a job on whichever cloud and region currently has the capacity and the best price, and is oriented around jobs and experiments. Cast AI is a managed platform that sources capacity for Kubernetes workloads and keeps managing them once running, including sharing, rightsizing and Spot handling. If the unit of work is a training job, SkyPilot fits naturally. If it is a long-running Kubernetes service, the managed path fits better.

Key takeaways

  • SkyPilot is an open-source job orchestrator. It runs workloads across 20+ clouds, selects the cheapest GPU available, and retries other regions automatically when capacity is unavailable.
  • Cast AI is a managed Kubernetes platform. Its OMNI Compute product extends existing EKS, GKE, and AKS clusters to source GPU capacity across AWS, GCP, and OCI without application code changes.
  • SkyPilot’s June 2026 SkyServe/inference endpoints release adds production inference serving with autoscaling, a gateway, TLS, and built-in metrics, making it a more complete platform than it was six months ago.
  • SkyPilot is the stronger choice for training experiments, batch jobs, and neocloud access (CoreWeave, RunPod, Lambda). Cast AI does not support these clouds.
  • Cast AI is the stronger choice for long-running Kubernetes inference services that need continuous cost optimization, automated GPU sharing, and per-team cost attribution.
  • Workload type determines the answer. Use SkyPilot for discrete jobs with a defined end. Use Cast AI for services that run indefinitely and need continuous management.

The short answer

If your team runs training experiments, hyperparameter sweeps, or batch jobs, especially on neoclouds like CoreWeave or RunPod, SkyPilot is the right tool. Its YAML-based interface is faster to start with than any Kubernetes setup, and its Managed Jobs handle spot preemptions and restarts automatically.

If your team runs long-running Kubernetes inference services on EKS, GKE, or AKS and wants continuous cost optimization without changing application code, Cast AI is the right tool. It sources GPU capacity from multiple clouds, manages Spot transitions with minimal disruption via PDB-aware eviction, and rightsizes nodes automatically.

Many teams use both. SkyPilot handles the training pipeline. Cast AI manages the serving layer. They operate in separate layers and do not conflict.

Comparison table

FeatureSkyPilot (OSS)Cast AI
Cloud reach20+ clouds incl. CoreWeave, RunPod, Lambda, Nebius, Vast.ai, and hyperscalersAWS, GCP, Azure, OCI (hyperscalers only)
Kubernetes supportYes, as one of many backends (EKS, GKE, AKS)Yes, native (EKS, GKE, AKS): core platform
GPU sharing automationNo: relies on existing cluster configurationYes: automates time-slicing (up to 48 replicas on A100 GPUs) and MIG (up to 7 instances on A100/H100)
Job managementYes (sky launch, sky jobs, Managed Jobs)No
Inference servingYes: SkyServe / inference endpoints (production-ready since June 2026)Yes: standard K8s Deployment + Cast AI autoscaler, optimized continuously
Spot recovery for jobsYes: restarts from checkpoint on preemptions, crashes, NCCL timeouts (user implements checkpointing)No
Cost optimization for servicesLimited: SkyServe has basic autoscalingYes: continuous rightsizing, Spot management, bin-packing
Continuous rightsizingNoYes
Managed vs self-hostedSelf-hosted (SkyPilot Platform: commercial managed tier, announced 2026)Fully managed
Multi-node distributed trainingYes: gang scheduling, checkpoint-based recovery on NCCL timeoutsNot suited (WAN latency across regions)
PricingFree (Apache 2.0)Quote-based (not public)
Primary use caseTraining jobs, batch workloads, and multi-cloud inference servingLong-running Kubernetes services: inference, serving, vector DBs

What SkyPilot does

SkyPilot is an Apache 2.0 open-source framework from the UC Berkeley Sky Computing Lab (docs.skypilot.ai, 2024). Ion Stoica, a SkyPilot co-founder, is also a co-founder of Databricks. SkyPilot supports 20+ clouds: AWS, GCP, Azure, OCI, CoreWeave, RunPod, Lambda, Nebius, Vast.ai, and others, including Kubernetes clusters and Slurm.

The core abstraction is the job. You define resources and a run command in YAML. SkyPilot provisions the cheapest matching instance, then retries other regions or clouds automatically when it hits a capacity error. That failover behavior is on by default, not an optional setting.

For distributed training, SkyPilot handles gang scheduling across multi-node clusters. Managed Jobs (sky jobs) adds fault tolerance: the system auto-recovers from node crashes, GPU failures, preemptions, and NCCL timeouts by restarting the job from the last saved checkpoint. Checkpointing must be implemented in your training code, SkyPilot does not insert checkpoint logic. For long training runs on spot instances, this recovery behavior reduces manual intervention significantly.

In June 2026, SkyPilot released its inference serving capabilities (SkyServe) for production use: autoscaling, a gateway, TLS, and built-in metrics (SkyPilot blog, “SkyPilot Endpoints: production-ready inference on every cluster you own,” June 23, 2026). Workloads run in your own cloud accounts and VPCs. As of 2026, SkyPilot reports 14M+ total downloads and 35% month-over-month GPU hours growth (source: SkyPilot Platform announcement blog, 2026; methodology not independently verified).

What Cast AI does

Cast AI is a managed Kubernetes platform. OMNI Compute, launched January 2026, extends existing EKS, GKE, and AKS clusters to source GPU capacity from any AWS, GCP, or OCI region. Application code does not change. Your team keeps its existing Kubernetes tooling.

Remote GPU nodes appear as virtual nodes inside your primary cluster through a Liqo-based multi-cluster architecture. The scheduler places pods normally. Cross-region provisioning happens below the Kubernetes API layer.

Continuous optimization runs after deployment: rightsizing, Spot selection, and bin-packing update automatically as workloads change. Per-workload and per-team cost attribution is built in via Kubernetes namespace and label tracking.

Cast AI is not a job orchestrator. It does not manage training job lifecycles, handle spot preemptions mid-run, or recover interrupted training experiments. For those needs, a separate tool is required.

Job-oriented vs service-oriented: the real dividing line

SkyPilot’s origins are in job orchestration. You define a task, it runs to completion, and the infrastructure disappears. This model fits training experiments, hyperparameter sweeps, and batch processing well.

SkyPilot’s June 2026 SkyServe release changes this picture. SkyServe adds production inference serving: autoscaling, a load-balancing gateway, TLS, and metrics. SkyPilot is now a more complete platform than it was six months ago, covering both discrete jobs and long-running serving.

Cast AI targets the workload that never stops: an inference API serving production traffic, a vector database with live queries, a model server handling user requests. These services need continuous management, not one-time provisioning decisions.

How inference deployment works on each side

The deployment models differ in a way that matters operationally. With SkyPilot, you deploy a serving job using sky serve up with a YAML spec. SkyServe manages the replica lifecycle, exposes an endpoint URL, and scales based on request volume. Infrastructure is provisioned and torn down by SkyPilot. You do not write Kubernetes manifests.

With Cast AI, you write a standard Kubernetes Deployment. The application layer is pure Kubernetes. Cast AI’s autoscaler manages what runs underneath: it selects the right node type, sources capacity across clouds, handles Spot transitions with minimal disruption via PDB-aware eviction, and rightsizes nodes as traffic patterns change. The serving team writes normal K8s specs. The infrastructure team enables Cast AI automation. The two roles stay separated.

The practical dividing line is still workload behavior. Jobs have a defined end; services run indefinitely. Cast AI continuously monitors resource consumption, rightsizes nodes, and manages Spot instances. SkyServe handles autoscaling for individual serving deployments, but Cast AI’s optimization layer covers the full Kubernetes cluster across all workloads simultaneously, not just a single inference deployment.

GPU utilization across Kubernetes clusters averages 5%, with Azure AKS at 2%, AWS EKS at 5%, and GCP GKE at 6% (Cast AI 2026 State of Kubernetes Optimization Report). A service running at 5% utilization is not a provisioning problem. It is an ongoing management problem.

Kubernetes integration on both sides

SkyPilot treats Kubernetes as one backend option among many. You can point sky launch at an existing EKS or GKE cluster, and SkyPilot schedules pods there (docs.skypilot.ai, 2024). It is not designed to manage long-running Kubernetes Deployments or StatefulSets. Its strength is job placement, not service lifecycle management.

Cast AI is Kubernetes-native by design. OMNI creates virtual nodes in your existing cluster. Remote GPU capacity from other regions or clouds appears as schedulable nodes in kubectl get nodes. Your Deployment spec does not change. Nothing at the application layer needs to know the workload is running on a node in a different region.

For teams already running on EKS, GKE, or AKS, Cast AI’s integration path requires no tooling changes. For teams that want to skip Kubernetes entirely and run jobs directly against cloud APIs, SkyPilot’s YAML-based interface is the faster path with significantly less infrastructure to manage.

GPU sharing and utilization

GPU sharing lets multiple workloads run on one physical GPU. This matters when inference replicas or batch jobs do not each require a full GPU’s resources and the math here is compelling. A 7B-parameter model running at FP16 precision requires roughly 14GB of VRAM (7 billion parameters × 2 bytes). A single A100-80GB card can host five such models natively before any sharing configuration, leaving the remaining ~10GB for KV cache overhead. With Cast AI’s automated time-slicing, the same A100 can serve substantially more replicas, up to 48 logical replicas on A100 GPUs, at the cost of shared GPU time rather than dedicated memory. For inference replicas with bursty, non-overlapping request patterns, this trades peak throughput headroom for a dramatic reduction in required node count.

SkyPilot has no native GPU sharing management. It provisions nodes and assigns workloads. Time-slicing and MIG configuration depend on whatever the underlying cluster already has configured before SkyPilot runs.

Cast AI automates GPU sharing configuration across GPU models with meaningfully different capabilities:

  • A10G (24GB): supports time-slicing and MPS only. MIG is not available on A10G at all — NVIDIA does not expose MIG partitions on this chip. Cast AI configures time-slicing for A10G workloads automatically.
  • L4 (24GB): supports MIG, but only the 1g.6gb profile — one isolated partition per physical GPU. This is a single-instance MIG configuration, distinct from A100/H100’s full seven-instance support. Useful for memory-isolated single replicas, not multi-tenant partitioning at scale.
  • A100 and H100: full MIG support with up to 7 isolated instances per GPU, each with dedicated memory and compute. Each instance is independently schedulable. Combining MIG partitions with additional time-slicing removes MIG’s memory isolation guarantee — use that configuration only when workloads do not require memory isolation.
  • Time-slicing (all supported GPUs): multiple workloads share one GPU’s compute time with no memory isolation. Cast AI supports up to 48 logical replicas per physical GPU on A100 GPUs in this mode.

For inference services where model size is smaller than full GPU memory, GPU sharing directly reduces hardware spend. Given average Kubernetes GPU utilization at 5%, most teams have substantial headroom to improve before they need additional hardware.

Operational burden: self-managed vs managed

SkyPilot is self-managed. You run and maintain the SkyPilot API server, handle upgrades, monitoring, and failure modes. SkyPilot Platform, a commercial managed offering backed by $20M in seed funding, was announced in 2026 (SkyPilot Platform announcement blog, 2026) and aims to reduce this overhead. Teams evaluating the managed tier should review the announcement blog directly for current availability and SLA commitments, as it was a recent launch at time of writing.

Cast AI is fully managed. The control plane, optimization logic, and OMNI networking run as a service. You connect your cluster, enable Phase 2 automation (Cast AI’s Phase 2 means the autoscaler is in active management mode, as opposed to Phase 1 read-only monitoring), and Cast AI handles provisioning decisions from that point. Phase 2 is required, teams on Phase 1 see recommendations but no automatic node provisioning or spot management.

The trade-off is control versus engineering time. Self-managed means deep configurability. Managed means you trade some configuration flexibility for operational hours your team does not spend on infrastructure.

Cost model

SkyPilot is free under Apache 2.0. You pay for the cloud compute it provisions. Its cost advantage for training comes from two mechanisms: spot GPU instances and cross-cloud arbitrage. Spot GPU instances typically save 60–80% versus on-demand pricing, depending on cloud, region, and GPU type. If AWS spot A100s are unavailable or expensive, SkyPilot routes to GCP or a neocloud automatically. Fewer than 2% of GPU workloads currently run on spot instances (Cast AI 2026 report), meaning most teams have not captured that savings potential. SkyPilot makes spot practical by automating the preemption handling and checkpoint-based restarts that otherwise make spot unreliable for training.

Cast AI pricing is quote-based. The platform provides continuous optimization: automatic spot selection, rightsizing, and cross-region sourcing. Teams provisioning continuously in cost-favorable regions save 2–5x versus on-demand rates for equivalent GPU capacity (Cast AI 2026 customer data). For context, AWS H200 Capacity Block prices rose 15% in January 2026, making active cost management more valuable as market prices shift (Cast AI GPU Price Report, 2026).

Neither cost model is inherently better. SkyPilot’s total cost is compute plus your operational time. Cast AI’s total cost is a platform fee plus compute, offset by what continuous optimization saves. The right comparison requires knowing your actual GPU utilization baseline.

Where SkyPilot is the stronger choice

SkyPilot is the right tool in these specific situations, without qualification:

Neocloud access

Cast AI covers AWS, GCP, Azure, and OCI. SkyPilot reaches CoreWeave, RunPod, Lambda, Nebius, Vast.ai, and 15+ additional clouds. If your team needs H100s on CoreWeave today because hyperscaler capacity is unavailable, SkyPilot gets you there. Cast AI does not.

Multi-node distributed training

SkyPilot’s gang scheduling and checkpoint-based recovery for NCCL timeouts are built for multi-node training. OMNI’s cross-cloud architecture crosses WAN boundaries, adding latency that is incompatible with tight allreduce communication patterns. For cross-cloud distributed training at scale, SkyPilot is the correct choice.

Teams avoiding Kubernetes

SkyPilot’s YAML interface is simpler than any Kubernetes configuration. Cast AI requires a Kubernetes cluster with Phase 2 activation (defined above). If your ML team does not operate Kubernetes and does not want to, SkyPilot removes that dependency entirely.

Training experiments and batch jobs

Hyperparameter sweeps, offline evaluation runs, and dataset preprocessing map directly to SkyPilot’s job model. Getting started takes a pip install and a short YAML file:

pip install "skypilot[all]"

A minimal SkyPilot YAML to launch a spot GPU training job:

# train.yaml — SkyPilot spot training job
name: my-training-run

resources:
  accelerators: A100:1
  use_spot: true

run: |
  python train.py --epochs 10 --checkpoint-dir /outputs

file_mounts:
  /outputs: s3://your-bucket/checkpoints

Run with sky launch -c my-cluster train.yaml. SkyPilot selects the cheapest available A100 across your configured clouds and retries automatically if the first region has no capacity.

The equivalent Cast AI entry point is a node template that defines what GPU capacity to source. A minimal OMNI node template for GPU inference:

# Cast AI OMNI node template — GPU inference (spot)
apiVersion: cast.ai/v1beta1
kind: NodeTemplate
metadata:
  name: gpu-inference-spot
spec:
  nodeConstraints:
    minCpu: 8
    maxCpu: 32
    minMemory: 32768   # MiB
    maxMemory: 131072
  gpuEnabled: true
  spot: true
  spotInterruptionHandling: true
  customLabels:
    workload: inference
    cast.ai/managed: "true"

SkyPilot provisions and terminates the node itself. Cast AI inserts the node into your existing Kubernetes cluster as a schedulable virtual node, your existing Deployment specs target it via standard node selectors or labels.

Spot-backed training with interruption recovery

SkyPilot Managed Jobs restart interrupted training from the last saved checkpoint when a preemption or NCCL timeout occurs. Recovery means restarting the job from checkpoint, not resuming an in-flight allreduce operation. Checkpointing logic must be implemented in your training code. A run without checkpoint saves restarts from the beginning.

Multi-cloud LLM inference serving

SkyServe (June 2026) adds production inference serving with autoscaling, a gateway, TLS, and metrics. For teams wanting multi-cloud LLM inference without Kubernetes operational overhead, SkyServe is worth evaluating directly against Cast AI’s serving capabilities before committing to either.

Where Cast AI is the stronger choice

Cast AI is the right tool in these specific situations:

Inference services with sustained traffic

Long-running inference APIs need continuous management: Spot handling, bin-packing, and rightsizing as request patterns shift over days and weeks. Cast AI monitors and adjusts the full cluster continuously — not just at deploy time. SkyServe handles autoscaling for individual serving deployments, but Cast AI’s optimization layer covers the entire Kubernetes cluster across all workloads running simultaneously. For a platform team managing 20 or more inference services, this eliminates a recurring class of Spot-related on-call incidents that otherwise surface at 3am.

Teams already on EKS, GKE, or AKS

Cast AI installs into your existing cluster. Application manifests, services, and ingress configuration remain unchanged. GPU nodes from additional clouds appear as standard Kubernetes nodes. The concrete first step is the OMNI docs page at docs.cast.ai/docs/omni-overview, which covers cluster connection and the architecture. Phase 2 activation is a toggle in the Cast AI console, not a code change, and must be enabled for Cast AI to provision nodes automatically rather than just surface recommendations. For teams that have run Kubernetes for years, the time from cluster connect to active cross-cloud GPU provisioning is typically measured in hours, not sprint cycles.

Multi-model serving with GPU sharing

When you run many model replicas and most do not each need a full GPU, Cast AI’s automated GPU sharing configuration reduces node count directly. Consider a cluster serving ten 7B-parameter models: each model needs ~14GB VRAM, so a single A100-80GB card handles five natively before any sharing. Cast AI’s automated time-slicing and MIG configuration, across A100, H100, and L4 GPUs, handles this without manual platform team intervention. For a platform team maintaining dozens of model versions simultaneously, the elimination of manual GPU sharing configuration at scale is the real operational win.

Per-team cost attribution and chargeback

Cast AI tracks GPU spend by Kubernetes namespace and label. For organizations running chargeback reports by team or product line, this is built into the platform. Building equivalent attribution on top of SkyPilot requires external tooling and manual instrumentation. For a FinOps team fielding monthly chargeback queries from ten or more internal teams, having namespace-level GPU cost data available without custom pipeline work removes a class of toil that compounds with every new model deployment.

Spot management for production services

Fewer than 2% of GPU workloads currently run on Spot (Cast AI 2026 report), primarily because Spot interruptions on services, as opposed to batch jobs, are operationally expensive to handle manually. Cast AI’s PDB-aware eviction handling makes Spot viable for long-running inference services by draining pods gracefully and rescheduling to available capacity before the Spot instance terminates. For a team running a production inference API on eight GPU nodes, shifting those nodes to Spot under Cast AI management can reduce monthly GPU spend by 60–70% without changing the application or accepting unplanned downtime.

Conclusion

SkyPilot and Cast AI address different problems at different layers. The decision is not which tool is better. It is which tool fits the workload shape.

SkyPilot handles GPU access and job orchestration. For training experiments across 20+ clouds including neoclouds, its cross-cloud spot arbitrage and job management are difficult to replicate with Kubernetes-native tools alone. Its June 2026 SkyServe release extends this into production inference serving, making it a credible platform for teams who want to avoid Kubernetes overhead entirely.

Cast AI handles continuous Kubernetes infrastructure management. For inference services on EKS, GKE, or AKS, it sources GPU capacity automatically from multiple clouds without application changes and optimizes continuously as workloads evolve. The GPU utilization data (5% average across clusters) makes the case for continuous management better than any benchmark: the problem is ongoing, not one-time.

A model fine-tuned on SkyPilot can be deployed to a Cast AI-managed inference endpoint using the same Kubernetes cluster. The two tools operate in separate layers. Many teams run SkyPilot for training and Cast AI for serving without conflict.

If you run inference workloads on EKS, GKE, or AKS and want to see what continuous GPU optimization looks like in practice, start with the OMNI overview or explore the best GPU optimization tools for Kubernetes AI.

Frequently Asked Questions

What is SkyPilot?

SkyPilot is an open-source framework (Apache 2.0) from the UC Berkeley Sky Computing Lab. It runs compute jobs across 20+ clouds and Kubernetes clusters, automatically selecting the cheapest available GPU instance and failing over to other regions or clouds when capacity is unavailable. It handles training jobs, batch workloads, and experiment pipelines. Since June 2026, it also supports production inference serving via SkyServe. See docs.skypilot.ai for full documentation.

What is the difference between Cast AI and SkyPilot?

SkyPilot is a job orchestrator: you submit a job, it finds the cheapest GPU across clouds and runs it to completion. Cast AI is a managed Kubernetes platform: it continuously optimizes long-running services by rightsizing nodes, managing Spot instances with PDB-aware eviction, and sourcing capacity across AWS, GCP, and OCI. SkyPilot is purpose-built for training jobs and batch workloads. Cast AI is purpose-built for production inference services on Kubernetes that need continuous management beyond initial provisioning. Many teams use both.

Does SkyPilot work with Kubernetes?

Yes. SkyPilot supports Kubernetes as one of its backends. You can point sky launch at an existing EKS, GKE, or AKS cluster, and SkyPilot will schedule pods there. It also supports SkyServe for inference serving on Kubernetes. SkyPilot treats Kubernetes as one option among many — it does not require Kubernetes and does not manage long-running Kubernetes Deployments. See docs.skypilot.ai for Kubernetes setup details.

Is SkyPilot free?

Yes. SkyPilot’s open-source version is free under the Apache 2.0 license. You pay only for the cloud compute it provisions, there are no license fees. SkyPilot Platform, a commercial managed offering announced in 2026, has separate pricing. The OSS version remains free.

Which is better for training jobs?

SkyPilot. It is purpose-built for training: gang scheduling for multi-node runs, automatic spot preemption recovery via Managed Jobs, and access to neoclouds like CoreWeave and RunPod that Cast AI does not support. One important caveat: Managed Jobs recovery means restarting the job from the last saved checkpoint – not resuming mid-computation. Checkpointing logic must be implemented in your training code. A run without checkpoint saves restarts from scratch on preemption.

How do I run GPU workloads across clouds?

The approach depends on your workload type. For training jobs: install SkyPilot with pip install "skypilot[all]", define your resources in a YAML file, and run sky launch. SkyPilot handles cloud selection and failover automatically. For long-running Kubernetes inference services: Cast AI OMNI Compute extends your existing EKS, GKE, or AKS cluster to source GPU capacity from other regions and clouds without changing your application. Start at docs.cast.ai/docs/omni-overview for the architecture overview and cluster connection steps. Phase 2 activation is a toggle in the Cast AI console — no code changes required — and enables automatic node provisioning across clouds.

Cast AI › Blog › Cast AI vs SkyPilot: Two Approaches to Finding GPUs Across Clouds