| Compute Savings PlansAWS | A $/hour spend rate, not a specific instance shapeApplies automatically across EC2 (any family/size/OS/region), Fargate, and Lambda | 1 or 3 years; All/Partial/No Upfront | Up to 66% off On-Demand, published | A footprint whose instance mix or region changes over timeWatch out: The most flexible option also earns the smallest discount of the three EC2 commitment types | AWS docs → |
|---|
| EC2 Instance Savings PlansAWS | A $/hour spend rate within one instance family in one regionSize and OS flex freely within the committed family/region; can't cross families or regions | 1 or 3 years; All/Partial/No Upfront | Up to 72% off On-Demand, published | A stable instance family in a stable region, with sizes still expected to changeWatch out: A planned migration off the committed instance family strands the whole commitment | AWS docs → |
|---|
| Standard Reserved InstancesAWS | A specific instance family, size, and Availability Zone or regionNone; it cannot be exchanged for a different configuration | 1 or 3 years; All/Partial/No Upfront | Up to 72% off On-Demand (AWS-published averages: ~40% at 1 year, ~60% at 3 years) | A genuinely fixed, predictable workload you're not planning to touchWatch out: Zero flexibility means the deepest discount of the three EC2 options, and the easiest to strand | AWS docs → |
|---|
| Convertible Reserved InstancesAWS | An instance family and region, but exchangeable for another of equal or greater valueCan be exchanged mid-term for a different family/size, for a slightly smaller discount than Standard RIs | 1 or 3 years; All/Partial/No Upfront | Up to 66% off On-Demand (AWS-published averages: ~31% at 1 year, ~54% at 3 years) | A workload you expect to reshape, where you still want an RI-style discountWatch out: The exchange right isn't free: it costs several points of discount versus a Standard RI | AWS docs → |
|---|
| Spot InstancesAWS | Nothing: no commitment at all, in either directionCancel anytime; AWS can also reclaim capacity with as little as ~2 minutes' notice | None: priced and billed like On-Demand, per second | Up to 90% off On-Demand, published | Stateless, interruption-tolerant work: batch jobs, CI runners, fault-tolerant worker fleetsWatch out: Not a substitute for a commitment decision; it's a different risk trade entirely (capacity risk instead of usage risk) | AWS docs → |
|---|
| RDS Reserved InstancesAWS | A specific database engine, instance class, and Deployment (single-AZ / Multi-AZ)Can modify within some limits; no cross-engine exchange | 1 or 3 years; All/Partial/No Upfront | Up to 69% off On-Demand in steady state (AWS-published: up to 30% No Upfront, up to 63% All Upfront at 3 years) | A database tier that runs the same shape continuouslyWatch out: ElastiCache and OpenSearch offer the same reserved-node model, so check each service's own pricing page before assuming RDS's rate carries over | AWS docs → |
|---|
| Resource-based Committed Use DiscountsGoogle Cloud | A specific vCPU/memory shape by machine seriesNone: tied to the committed machine series | 1 or 3 years | Up to 55% for most machine series, up to 70% for memory-optimized series, published | A stable machine-series footprint you're confident won't moveWatch out: Discount varies materially by machine series; the headline number isn't the same for every shape | Google Cloud docs → |
|---|
| Flexible (spend-based) Committed Use DiscountsGoogle Cloud | A $/hour spend rate, applied automatically across eligible machine types in a regionFlexes across machine series/size within the commitment, similar in spirit to an AWS Compute Savings Plan | 1 or 3 years | ~28% (1-year) to ~46% (3-year) for general-purpose machines, published; varies by series | A footprint whose machine mix shifts, where you still want automatic coverageWatch out: Meaningfully lower discount than resource-based CUDs; flexibility again costs discount depth | Google Cloud docs → |
|---|
| Reserved VM InstancesAzure | A specific VM size/series in a specific regionExchanges and cancellations are possible subject to Azure's reservation exchange policy | 1 or 3 years | Up to 72% vs. pay-as-you-go, published (Microsoft's own footnote: actual savings vary by location, instance type, and usage) | A fixed VM footprint in a region you're not planning to leaveWatch out: Microsoft's headline figure is a specific scenario, not a guaranteed rate for every VM/region combination | Azure docs → |
|---|
| Savings Plan for ComputeAzure | A $/hour spend rate, applied automatically across eligible compute servicesFlexes across VM series, size, region, and OS, plus some container and App Service usage | 1 or 3 years | Up to 65% vs. pay-as-you-go, published (Microsoft's own footnote: real-world savings observed between 11% and 65%) | A shifting compute footprint where automatic, hands-off coverage matters more than the deepest possible rateWatch out: The 11–65% real-world range is wide, so don't plan a budget off the 65% headline alone | Azure docs → |
|---|
| Provisioned ThroughputAmazon Bedrock | A guaranteed inference capacity level (model units/hour) instead of a per-token rateNone once purchased for the term; a no-commitment hourly rate also exists as a fallback, at a premium | 1 month or 6 months (per-model; some model providers require contacting AWS directly) | 6-month commitment prices meaningfully below 1-month for the model families that publish self-serve rates. It varies by model, with no single site-wide percentage | Latency-sensitive, high and predictable-volume inference where on-demand token pricing would cost more at that volumeWatch out: This buys guaranteed capacity, not a discount on token price. Model it against your actual token volume, not against the sticker discount | Amazon Bedrock docs → |
|---|
| Message Batches APIAnthropic | Nothing: no term, no minimum volume, an asynchronous processing mode instead of a real-time callUse it request-by-request; no ongoing obligation of any kind | None: per-request, asynchronous, typically completes within 24 hours | 50% off both input and output tokens vs. standard synchronous pricing, published | Anything that doesn't need a synchronous response: evals, backfills, bulk classification, offline scoringWatch out: No SLA for when a batch completes within the window, so it's not a fit for anything latency-sensitive | Anthropic docs → |
|---|
| Batch APIOpenAI | Nothing: same commitment-free shape as Anthropic's batch modeUse it request-by-request; no ongoing obligation of any kind | None: per-request, asynchronous, completes within 24 hours | 50% cost discount vs. synchronous API pricing, published | Anything that doesn't need a synchronous responseWatch out: Same latency trade as Anthropic's batch mode, so don't route anything user-facing through it | OpenAI docs → |
|---|
| Prompt cachingAnthropic | Nothing: no term, no commitment; a cache write/read mechanic, not a purchased discountOpt in per-request with a single field; no ongoing obligation | Cache entries live 5 minutes or 1 hour per write, refreshed on each hit | Cache reads cost 0.1x base input price; writes cost 1.25x (5-min) or 2x (1-hour) base input price, published | Repeated system prompts, long reference documents, or conversation history reused across callsWatch out: This is the AI-side analogue of a commitment discount that carries zero downside risk. There's no version of this page's break-even math for it, because there's nothing to strand | Anthropic docs → |
|---|