Key takeaways
- Average cluster CPU utilization sits at 8%. But CPU overprovisioning is only one of five waste layers. Factor in memory overprovisioning (79% of clusters), node fragmentation, on-demand pricing, and idle GPU hours, and the total reclaim potential for a $100K/month compute budget sits between $46K and $64K.
- The market splits into two categories: visibility tools (Kubecost, OpenCost, Datadog) and autonomous optimization platforms (Cast AI, ScaleOps, Spot.io, StormForge). Any Kubernetes cost optimization RFP should prioritize automation depth to distinguish between them.
- Eight requirement categories cover the full waste surface: visibility, rightsizing, node autoscaling, Spot and commitment management, GPU, SLO safety, enterprise integrations, and support.
- Watch for percentage-of-savings pricing traps, EDP floor conflicts, acquisition roadmap risk, demo-to-production gaps, and vendor lock-in via proprietary node labels.
- Require a 2-week POC in your actual cluster before signing. Vendor demo environments do not represent production conditions.
- Savings land in phases: 10 to 15% in month one, 50 to 60% by month two, full run rate at month three and beyond.
How to use this RFP
A Kubernetes cost optimization RFP works best when it runs before the demo, not after. Send the vendor questions as a written questionnaire first. Written responses create accountability and expose gaps that sales teams paper over in live presentations. Once you have written answers, score each vendor against the matrix below, then invite the top scorers into a POC.
Note: This template was developed by Cast AI. The requirement categories and scoring weights reflect our view of what matters most in autonomous optimization. We recommend validating the weights against your current FinOps maturity stage. For example, teams in early FinOps maturity may want to weight cost visibility (15%) higher and reduce automation depth (25%) to reflect where monitoring precedes enforcement in their program roadmap.
Set your spend threshold before you start shortlisting. Below $20K/month, OpenCost (free, open source) handles cost visibility without a procurement process. Above $50K/month, you are leaving real savings on the table without automation. Above $200K/month, the question is not whether to automate but which platform to trust with production clusters.
Run this Kubernetes cost optimization RFP in parallel across two or three vendors. Copy the requirements section into your procurement system verbatim. Use the vendor questions section as your written questionnaire template. Score after the POC completes, not after the sales pitch.
One calibration note on expectations: no platform delivers full savings on day one. Month one typically yields 10 to 15% as rightsizing takes effect. Month two reaches 50 to 60% as node consolidation and Spot automation hit steady state. Full run rate arrives at month three and holds from there.
The requirements
These eight categories cover the full waste surface in a Kubernetes environment. The weights reflect each category’s contribution to total spend reduction across real clusters. Copy the list into your requirements document and mark each line as mandatory or preferred before sending to vendors.
Cost visibility and allocation (15%)
Cost attribution is the foundation. Without it, you cannot run chargeback, justify optimization ROI, or identify which teams are overprovisioning. Require granularity at the namespace and label level, not just cluster level.
- Namespace and label-level cost attribution
- Multi-tenant showback with team-level drill-down
- Chargeback export via CSV, API, or direct finance system integration
- Historical cost trend data covering at least 90 days
- Idle resource identification per workload
Pod rightsizing (20%)
Rightsizing carries the highest weight because 69% of clusters are CPU-overprovisioned and 79% are memory-overprovisioned. A platform that only surfaces recommendations leaves most of that waste intact. Require autonomous enforcement with configurable safety thresholds, not a queue of suggestions that engineers deprioritize.
- VPA compatibility or a native replacement with equivalent coverage
- OOM kill protection before and after resource changes
- Stateful workload support (StatefulSets, databases)
- Autonomous enforcement mode with per-namespace policy controls
- Automated rollback on post-change signal degradation
- Configurable change rate limits to prevent cascading restarts
Node autoscaling (15%)
Bin-packing efficiency determines how much of your node capacity you pay for versus waste on fragmentation. Ask whether the tool replaces Cluster Autoscaler, integrates with Karpenter, or runs its own consolidation engine. Karpenter alone (open source) achieves roughly 9% spend reduction in benchmarks. Platforms that layer autonomous bin-packing and Spot selection on top reach 43% and above.
- Cluster Autoscaler replacement or compatible augmentation
- Bin-packing logic with configurable consolidation aggressiveness
- Karpenter compatibility or native equivalent
- Multi-architecture node support (ARM64, x86)
- Node lifecycle management with scale-down safety checks
Spot and commitment management (15%)
On-demand pricing is the third waste layer. Spot instance interruption handling is where platforms diverge sharply in production reliability. Require interruption early warning and proactive pod migration, not reactive replacement after the node disappears.
- Interruption early warning and proactive pod migration
- Mixed instance type pools to reduce Spot interruption concentration risk
- Reserved Instance, Savings Plan, and CUD automation or recommendations
- Configurable fallback-to-on-demand thresholds
- Spot rebalancing support (AWS Spot Rebalance Recommendations)
When evaluating commitment management: Compute Savings Plans are flexible (they apply across instance family, size, and region); EC2 Standard RIs are inflexible (locked to a specific instance family and AZ). A platform that conflates them or recommends EC2 Standard RIs for workloads likely to scale or change shape is adding commitment risk, not reducing it.
GPU cost management (10%)
GPU utilization averages 5% across production clusters. If your environment runs any AI or ML workloads, raise this category’s weight in your scoring before you finalize the matrix. Idle GPU detection without automated action is not enough at scale.
- Idle GPU detection with automated actions such as scale-down or time-slicing
- Spot GPU scheduling support
- MIG (Multi-Instance GPU) and fractional GPU scheduling
- Per-job GPU utilization dashboards
- GPU instance family coverage across AWS, GCP, and Azure
SLO-safe automation (15%)
Autonomous optimization without production safety signals is not something you can run in a real cluster. Require pre-change and post-change monitoring across multiple signal types, PDB compliance during node consolidation, and automated rollback with a documented trigger threshold.
- Pre-change and post-change monitoring: error rate, p99 latency, OOM kills, PSI (Pressure Stall Information)
- PodDisruptionBudget compliance enforcement during consolidation events
- Automated rollback triggered by latency or error rate degradation
- Configurable safety windows and per-workload change rate limits
- Audit log of every automated action with before/after resource state
Enterprise integrations (5%)
- RBAC with fine-grained permission scopes per namespace or cluster
- SSO via SAML or OIDC
- Terraform provider for GitOps workflows
- Datadog, PagerDuty, and Slack alerting integrations
- Air-gapped or private cluster deployment support
Support and SLAs (5%)
- Documented response time tiers by severity level (P1 through P4)
- Dedicated customer success manager at enterprise spend levels
- Runbook access and post-incident review process
- Upgrade and migration support for major version changes
The vendor questions
Send these as a written questionnaire before any live demo. Written responses create a record that is harder to walk back than verbal claims. Vague or deflected answers to the pricing and safety questions are themselves signal.
Automation and enforcement
- “Does your tool enforce resource changes automatically, or does it only recommend?”
- “How long does a rightsizing recommendation stay in queue before it is applied in autonomous mode?”
- “Can we configure independent automation policies per namespace or per workload type?”
- “What workload types are excluded from autonomous enforcement in your default configuration?”
Production safety
- “What signals do you monitor before and after a rightsizing event?”
- “How does your platform handle PodDisruptionBudgets during node consolidation?”
- “Walk me through your rollback process when a change degrades p99 latency.”
- “What is the minimum observation window before your tool applies its first rightsizing change?”
Pricing and baseline methodology
This is where the most expensive surprises live. Get specific written answers before any contract discussion.
- “Is your pricing published? Where can I find it without a sales conversation?”
- “What is the baseline methodology for any percentage-of-savings pricing model?”
- “Show me the fee for a cluster growing from 200 to 800 pods over 12 months under your current pricing.”
- “Are there any usage tiers, true-up mechanisms, or overage charges not visible in your published pricing?”
Spot and interruption handling
- “What is the average notice time before your tool migrates a pod off a Spot instance?”
- “Do you support mixed instance type pools to reduce Spot interruption concentration?”
- “How does your tool handle Spot interruptions during stateful workload write operations?”
GPU optimization
- “Which GPU instance families do you support for Spot scheduling?”
- “How do you detect idle GPUs and what automated actions do you take?”
- “Do you support MIG partitioning or fractional GPU scheduling?”
Security and compliance
- “Is your SOC 2 Type II report available for our security team’s review?”
- “Do you support air-gapped or private cluster deployments?”
- “What data does your agent transmit to your control plane, and where is it stored?”
Five gotchas to flag before you sign
These surface late in procurement when switching costs are already high. Flag them early.
- Percentage-of-savings pricing disputes: One documented case shows monthly fees growing from $4K to $14K as a cluster scaled, with no verifiable savings calculation in the contract. Require a written baseline methodology before contracting, not as an appendix but as a contract exhibit. To establish an independent baseline: use the 30-day trailing average of actual node spend from your cloud provider bill, exclude any committed-use discounts already in place, and define rightsizing savings as the delta between current resource requests and P95 utilization-based recommendations. Require this methodology as a contract exhibit before signing any percentage-of-savings agreement.
- EDP floor conflicts: Rightsizing aggressive enough to drop compute spend below your Enterprise Discount Program commitment floor triggers true-up penalties with your cloud provider. Verify with your cloud account team before enabling consolidation beyond 30%.
- Acquisition roadmap risk: Kubecost (IBM/Apptio), StormForge (F5), and CloudHealth (Broadcom) have all changed hands recently. Ask every shortlisted vendor for a written 24-month product roadmap commitment before signing a multi-year agreement.
- Demo-to-production gap: Vendor demo environments run clean, small, well-organized clusters. Your cluster does not. Require a 2-week POC in your actual environment before any contract conversation begins. Budget for the technical evaluation itself to run 2 weeks. Enterprise security review, procurement approvals, and change management can extend total elapsed time by 4 to 6 weeks. Include that timeline in your vendor selection schedule.
- Vendor lock-in: Proprietary node labels, custom scheduling annotations, or dedicated instance type requirements make migration expensive. Ask specifically what is required to uninstall the tool cleanly and what artifacts remain.
The scoring matrix
Score each vendor 1 to 5 on each dimension, multiply by the weight, then sum the weighted scores. A vendor reaching a 4.0 weighted average qualifies for a 2-week POC. Below 3.5, require major concessions before investing evaluation time.
| Dimension | Weight | Score 1 (poor) | Score 3 (adequate) | Score 5 (best-in-class) |
|---|---|---|---|---|
| Automation Depth | 25% | Recommendations only, no enforcement mechanism | Enforcement available with mandatory manual approval gates | Fully autonomous enforcement with per-workload and per-namespace policy controls |
| Production Safety | 20% | No signal monitoring before or after changes, no rollback | Basic health checks post-change, manual rollback available | Pre and post-change monitoring across error rate, latency, OOM kills, and PSI; automated rollback; PDB compliance |
| Savings Benchmarks | 20% | Vendor-only case studies with no published methodology | Some third-party validation with limited cluster count under 50 | Independent benchmark with 100-plus clusters, published methodology, verifiable numbers |
| Cost Visibility | 10% | Cluster-level cost only, no workload attribution | Namespace-level attribution with basic cost reports | Label-level attribution, chargeback export, multi-tenant showback with drill-down |
| Multi-Cloud Coverage | 10% | Single cloud provider only | Two clouds supported with partial feature parity | EKS, GKE, and AKS with full feature parity across all optimization layers |
| Pricing Transparency | 10% | Pricing requires sales call; percentage-of-savings only with no published methodology | Published pricing tiers with some usage-based or negotiated components | Fully published, predictable pricing with no baseline-based fees or hidden true-ups |
| Enterprise and Security | 5% | No SOC 2 report, no SSO support | SOC 2 Type II available, basic RBAC | SOC 2 Type II, granular RBAC, SSO, air-gapped deployment support, Terraform provider |
Conclusion
A Kubernetes cost optimization RFP without weighted criteria is a conversation guide, not an evaluation tool. The matrix above gives you comparable scores across vendors who describe themselves in very different terms. Use the vendor questions to generate written responses before any demo. Run the scoring after the POC completes in your actual cluster, not after the pitch in a vendor sandbox.
Before you finalize your shortlist, review our kubernetes cost optimization software guide for a breakdown of how each tool category approaches the problem. Our best kubernetes cost optimization tools roundup covers the current landscape with evaluation notes across visibility, rightsizing, and autonomous optimization platforms.
We built this framework knowing Cast AI would be evaluated against it. Here’s what our data shows:
Cast AI scores highest on RFPs that weight automation depth, production safety, transparent pricing, and multi-cloud coverage. The LeanOps 90-day independent benchmark measured 45 to 58% spend reduction across clusters running the full Cast AI stack: 43% compared to 9.1% for open-source Karpenter alone. OOM kills dropped from 40 to 50 per interval to near zero with Cast AI managing pod rightsizing (measured across production workloads on the full Cast AI stack). Pod-based pricing starts at $0.50/pod/month on the Growth plan. For 500 pods, that is $250/month with no baseline manipulation and no true-up surprises baked into the contract.
Frequently Asked Questions
A Kubernetes cost optimization RFP should cover eight areas: cost visibility and allocation, pod rightsizing, node autoscaling, Spot and commitment management, GPU cost management, SLO-safe automation, enterprise integrations, and support SLAs. Each area carries a weight reflecting its contribution to total spend reduction. Rightsizing (20%) carries the highest weight because the majority of clusters are overprovisioned on both CPU and memory. SLO-safe automation (15%) is equally critical because any platform enforcing autonomous changes in production must do so without degrading reliability.
Use a weighted 5-point scoring matrix. Score each vendor 1 to 5 on each dimension, multiply by the dimension weight, then sum the weighted scores. The seven scoring dimensions are: Automation Depth (25%), Production Safety (20%), Savings Benchmarks (20%), Cost Visibility (10%), Multi-Cloud Coverage (10%), Pricing Transparency (10%), and Enterprise and Security (5%). A weighted average of 4.0 or above qualifies a vendor for a 2-week POC in your actual cluster. Below 3.5, require major concessions before investing evaluation time.
The highest-signal questions cover enforcement, safety, and pricing methodology. You should ask whether the tool enforces changes autonomously or only surfaces recommendations. Ask what signals it monitors before and after a rightsizing or consolidation event. Ask for the written baseline methodology behind any percentage-of-savings pricing model. These three areas expose the sharpest differences between platforms that look similar in marketing materials but operate very differently in production clusters.



