Every cloud and AI price change that moved a bill, sourced straight from the provider that made it, cuts and increases alike, with the concrete move to make in response.
Every entry below comes from the provider's own announcement, blog post, or pricing page — linked at the end of each one — verified as of 2026-08-04. No third-party recaps, no aggregators. This log tracks cuts as well as increases, because a log that only shows increases is marketing, not a record. See the full tools hub and resources for the rest of what CostMon publishes.
30price changes tracked
12price cuts
5price increases
10providers covered
2023–2026date range
Filter the price change log
Provider
Type
JavaScript is off, so every entry below is shown — filtering by provider or type needs it enabled.
2026
OpenAICut
OpenAI cut GPT-5.6 Luna's price 80% and Terra's 20%, leaving the flagship tier untouched.
Luna −80%, Terra −20%
What changed. Citing inference-stack efficiency gains, OpenAI dropped Terra to $2/$12 per million tokens and Luna to $0.20/$1.20, while the top-tier Sol model's price stayed flat.
Who it hits. High-volume, latency-tolerant workloads on Luna or Terra: classification, bulk summarization, agent sub-tasks.
What to do
Confirm the lower rate appears on your next invoice, then check whether tasks still pinned to Sol can safely move down to the now much cheaper Luna or Terra.
GitHub replaced Copilot's flat premium-request billing with usage-based AI credits.
$10-$39 credits/mo
What changed. Every Copilot plan now gets a monthly allotment of GitHub AI Credits consumed by actual token usage, instead of a flat per-request charge that billed a quick question the same as a multi-hour agent session.
Who it hits. Teams running heavy autonomous coding sessions, previously undercharged, and light users, previously overcharged, under the old flat-request model.
What to do
Track per-seat credit burn from June 1, 2026, and budget for the promotional credit boost on Business/Enterprise expiring after August 2026.
GPT-5.5 launched priced higher than GPT-5.4, betting token efficiency would offset it.
$5 / $30 per M tokens
What changed. GPT-5.5 launched at roughly double GPT-5.4's per-token rate. OpenAI's own announcement frames this as a deliberate increase, offset by claimed gains in token efficiency per task.
Who it hits. High-volume agentic or coding workloads that upgrade by default and inherit the new rate without checking whether the efficiency claim holds for their traffic.
What to do
Run a side-by-side cost comparison on real workloads before moving production traffic. Confirm the token savings offset the higher headline price for you.
Vercel cut Turbo build machine pricing 16% and unified every tier to one per-CPU rate.
−16% ($0.126→$0.105/min)
What changed. All build machine tiers now bill at a flat $0.0035 per CPU per minute, dropping the 30-CPU Turbo machine from $0.126 to $0.105 a minute.
Who it hits. Pro and Enterprise teams defaulted onto Turbo build machines and billed by build-CPU-minutes.
What to do
Confirm the lower rate applied on your next invoice, and re-check whether Turbo is still worth it over Standard or Enhanced now that the price gap has narrowed.
VPC Encryption Controls stopped being a free preview and started billing per VPC-hour.
$0.15 / VPC-hour (us-east-1)
What changed. After a free preview that ran from November 2025, AWS began charging a fixed hourly rate for every non-empty VPC running Encryption Controls in monitor or enforce mode. Empty VPCs stay free. The announcement itself doesn't print a rate. AWS's VPC pricing page puts it at $0.15 per hour in US East (N. Virginia), rising to $0.31 in the priciest regions.
Who it hits. Anyone who turned Encryption Controls on broadly during the free preview across many VPCs and never revisited it.
What to do
Audit which VPCs need enforce mode versus monitor mode, and turn it off on non-empty VPCs where it isn't earning its per-hour charge. Check your own region's rate too; it's up to twice the US East price.
Gemini 3 Flash turned on context caching by default, cutting repeated-token cost up to 90%.
up to −90% w/ caching
What changed. Gemini 3 Flash launched at $0.50 per million input tokens and $3 per million output tokens, with context caching standard rather than opt-in, plus a 50% discount for asynchronous Batch API jobs.
Who it hits. Repeated-context workloads (agents with long system prompts, RAG over the same documents, multi-turn chat), plus any non-interactive bulk job.
What to do
Confirm caching is active on your repeated prompts, and route non-interactive jobs through the Batch API to capture the 50% discount.
Snowflake moved Snowpipe to one flat per-GB rate, cutting ingestion cost for some teams over 50%.
flat 0.0037 credits/GB
What changed. Snowpipe pricing, previously based on compute consumed and file count, moved to a single 0.0037-credits-per-GB rate across file ingestion and streaming.
Who it hits. High-frequency, small-file ingestion pipelines that used to pay file-count-driven overhead.
What to do
Check ingestion cost after the automatic switchover, and stop over-batching files purely to dodge per-file charges. Billing is now per-GB regardless of file count.
Claude Opus 4.5 cut Opus-tier pricing to $5/$25 per million tokens.
$5 / $25 per M tokens
What changed. Opus 4.5 launched at $5 input / $25 output per million tokens, a cut Anthropic frames as putting frontier-tier capability within reach of more teams.
Who it hits. Teams that avoided the Opus tier on cost grounds for complex agents, deep research, or hard coding tasks.
What to do
Re-evaluate workloads kept on Sonnet purely for cost. At the new rate, Opus may be affordable for tasks where the extra capability wasn't worth the old premium.
Gemini 3 Pro launched as Google's new flagship, at a higher output rate than 2.5 Pro.
$2 / $12 per M tokens
What changed. Google's new flagship reasoning and agentic model launched at $2 per million input tokens and $12 per million output tokens for prompts up to 200K tokens.
Who it hits. Teams migrating agentic, coding, or reasoning workloads from Gemini 2.5 Pro or 1.5 Pro to the new flagship.
What to do
Benchmark real task cost against your current model before switching. The higher output rate needs to be earned back in fewer tokens per task, not assumed.
Cloud Build repriced its e2 machines by region, raising build minutes in some and trimming others.
e.g. $0.064→$0.0705/build-min (us-east1)
What changed. Per-build-minute pricing for e2 machine types moved from one flat rate to a region-by-region table: us-east1 (South Carolina) went up across e2-medium, e2-highcpu-8, and e2-highcpu-32, while us-central1 (Iowa) went down slightly on the same machine types.
Who it hits. CI/CD pipelines pinned to a region that landed on the wrong side of the new table.
What to do
Check your build pool's region against the new per-region rate table, and shift latency-insensitive builds to a cheaper region if yours got more expensive.
Claude Haiku 4.5 launched at a third of Sonnet 4's price for comparable coding performance.
$1 / $5 per M tokens
What changed. The new small model launched at a flat $1 input / $5 output per million tokens, positioned by Anthropic as Sonnet-4-level coding at a third of the cost.
Who it hits. Teams running Sonnet-tier models for routine coding or high-volume agent steps that don't need frontier reasoning.
What to do
Benchmark current Sonnet-4-class workloads against Haiku 4.5 and shift eligible traffic down a tier, checking quality on your own tasks first.
Azure cut Ultra Disk IOPS and throughput pricing in UK South by up to 80%.
−60% IOPS / −80% throughput
What changed. In UK South, provisioned IOPS on Ultra Disk dropped from $0.06205 to $0.02482 per month, and provisioned throughput from $0.40588 to $0.08103 per MBPS per month. Per-GiB capacity pricing didn't move.
Who it hits. IOPS- or throughput-heavy Ultra Disk workloads in UK South, like SAP HANA or high-transaction databases.
What to do
Re-run the disk cost estimate for UK South workloads at the new per-unit rates before assuming last quarter's numbers still hold.
GKE folded paid multi-cluster management into the Standard tier at no extra charge.
fleet mgmt now $0
What changed. As part of moving to a single paid GKE tier, Fleets, Teams, Config Management, and Policy Controller are now included with GKE Standard. All four were previously gated behind a pricier tier or add-on.
Who it hits. Teams running multiple GKE clusters that were paying extra just to get fleet management or policy enforcement across them.
What to do
Check your current GKE edition. If you upgraded purely for fleet management or Policy Controller, you can likely drop back to Standard and keep the feature.
GPT-5 replaced the GPT-4o/o-series lineup with one model family sold at three price points.
$1.25 / $10 per M tokens
What changed. GPT-5 launched at $1.25 per million input tokens and $10 per million output tokens, with mini and nano tiers underneath it: one family instead of separate reasoning and non-reasoning models.
Who it hits. Teams that were routing between GPT-4o and o-series reasoning models to manage cost, now facing one family to re-tier by task.
What to do
Map existing GPT-4o and o3 call sites to gpt-5, gpt-5-mini, or gpt-5-nano by task difficulty, and re-benchmark spend rather than assuming a like-for-like swap.
Vercel switched Functions billing to active CPU time instead of full wall-clock duration.
−53% ($0.318→$0.149/hr)
What changed. A Standard-size function running at 100% active CPU now costs about $0.149 an hour instead of $0.318, because idle time waiting on a database call or an LLM response no longer bills as CPU time.
Who it hits. AI inference and agent workloads with significant idle time waiting on a slow upstream call.
What to do
Re-run cost estimates for Functions-heavy AI workloads under the new model rather than assuming automatic savings, and confirm your plan has switched over.
Gemini 2.5 Flash's stable release raised input price but cut output price and merged two rate tiers into one.
in +100%, out −29%
What changed. Going stable, Google removed 2.5 Flash's separate thinking vs. non-thinking price tiers, raising input from $0.15 to $0.30 per million tokens while cutting output from $3.50 to $2.50.
Who it hits. Teams that budgeted around the preview's lower input price, or leaned on non-thinking mode's cheaper output rate.
What to do
Recompute cost-per-request at the new $0.30/$2.50 blended rate before migrating off the preview model. Input-heavy workloads got pricier even though output-heavy ones got cheaper.
AWS cut On-Demand GPU instance prices by up to 45%.
up to −45% (P5)
What changed. On-Demand pricing dropped across the P5, P5en, P4d, and P4de instance families: P5 by up to 45%, P5en by up to 26%, P4d/P4de by up to 33%.
Who it hits. ML teams training or serving on P5/P4d GPU instances who stayed on-demand instead of committing to a Savings Plan.
What to do
Compare the new On-Demand rate against your existing Savings Plan commitment. It may no longer be worth the lock-in.
Anthropic cut tool-use output tokens on Claude 3.7 Sonnet by up to 70%.
up to −70% output tokens
What changed. A new token-efficient tool-use mode cut output tokens consumed on tool-calling workloads, and prompt-cache reads stopped counting against input-token-per-minute rate limits.
Who it hits. Agentic and tool-calling workloads on Claude 3.7 Sonnet that were hitting rate limits or paying for verbose tool-call output.
What to do
Turn on token-efficient tool use for tool-calling workloads, and check whether old throttling logic built around the rate limit is still needed.
AWS cut GuardDuty's S3 malware-scanning price by 85%.
−85% ($0.60→$0.09/GB)
What changed. The per-GB price for the data-scanned dimension of Malware Protection for S3 dropped from $0.60 to $0.09 in US East (N. Virginia), applied automatically to existing customers.
Who it hits. Teams running malware scanning across high-volume S3 buckets, where the data-scanned line used to dominate the bill.
What to do
Re-run the cost projection for buckets you skipped scanning because it was too expensive. Coverage that didn't pencil out before likely does now.
MongoDB merged its Shared and Serverless tiers into one Flex tier with a capped bill.
$8 base, capped $30/mo
What changed. Atlas Flex bills an $8 base fee plus usage up to 500 ops/sec, capped at $30 a month. It replaces both the old Shared clusters and standalone Serverless instances.
Who it hits. Small teams and dev/test workloads previously on Shared or Serverless clusters.
What to do
Plan the migration off Shared or Serverless ahead of MongoDB's cutover, and model your peak burst usage against the $30 cap before assuming it still fits.
GitHub made Copilot free in VS Code, capped at 2,000 completions a month.
2,000 completions free/mo
What changed. Any GitHub account gets capped code completions and chat in VS Code: 2,000 completions and 50 chat messages a month, with no trial or subscription required.
Who it hits. Individual developers who previously had no free path into Copilot beyond a time-limited trial.
What to do
Point casual or evaluation users at the free tier instead of provisioning a paid seat, and watch for contributors who lean on it in place of a managed seat.
Cloud SQL began billing instances left running on end-of-life MySQL or PostgreSQL versions.
new per-vCPU/hr fee
What changed. Instances still running a MySQL or PostgreSQL major version past its community end-of-life get auto-enrolled in a paid "extended support" add-on, billed per vCPU-hour (or per instance-hour on shared-core). Google waived the charge through April 2025, then started billing it.
Who it hits. Anyone who put off a major-version upgrade and is still running the old version in production.
What to do
Check instance versions against Cloud SQL's end-of-life schedule and upgrade off any EOL major version instead of paying the ongoing surcharge.
Azure zeroed out the bandwidth charge for traffic between Availability Zones in the same region.
→ $0/GB
What changed. Microsoft eliminated the per-GB charge for data moving across Availability Zones within a region, for both private and public IP traffic.
Who it hits. Zone-redundant deployments: replicated databases, storage, and VM fleets that talk across zones for reliability.
What to do
Revisit any old cost-optimization advice that steered you away from zone-redundant architecture to dodge bandwidth charges. That tradeoff is gone.
AWS paired the new IPv4 charge with 750 free IPv4-hours a month, for new accounts only.
750 IPv4-hrs/mo free
What changed. New AWS accounts get 750 hours a month of public IPv4 usage at no charge when they launch an EC2 instance with a public address. That's enough to cover one always-on instance.
Who it hits. New accounts running a single small workload. Established accounts with many public IPs get nothing from this.
What to do
Don't assume this offsets your bill. Check whether your account qualifies as "new" before crediting it against the IPv4 charge.
Google Cloud raised the hourly charge on in-use external IPv4 addresses.
$0.004→$0.005/hr
What changed. The External IP Charge for a standard VM rose from $0.004 to $0.005 an hour, and for a Spot VM from $0.002 to $0.0025, for every in-use external IPv4 address, on the same schedule as AWS's own IPv4 pricing move.
Who it hits. VMs with a public address attached: bastions, NAT boxes, anything public-facing that didn't need to be.
What to do
Move workloads to internal-only addressing behind Cloud NAT or a load balancer, and keep external IPs reserved for the resources that genuinely need one.
AWS started charging for every public IPv4 address, attached or not.
+$0.005 / IP-hour
What changed. Every public IPv4 address across AWS started accruing an hourly charge, whether it's attached to a running instance or just sitting on an Elastic IP nobody released.
Who it hits. Fleets of small instances that each hold their own public IP, and any account with stale Elastic IPs it forgot to release.
What to do
Run VPC IP Address Manager's Public IP Insights, move workloads behind a shared NAT gateway or load balancer, and release unattached Elastic IPs.
BigQuery raised on-demand query pricing 25% the same day it retired flat-rate slots.
+25% on-demand
What changed. Google retired flat-rate annual, flat-rate monthly, and flex-slot commitments in favor of three new editions billed per-second, and raised the on-demand analysis price 25% across all regions on the same date.
Who it hits. Teams paying per-bytes-scanned for ad-hoc queries, and anyone on a flat-rate or flex-slot commitment who now has to migrate to an edition.
What to do
Cut bytes scanned with partitioning and clustering before the 25% increase compounds, and evaluate a slot-based edition if your on-demand spend is now significant.
Azure raised local-currency prices in the UK, EU, and Nordics to track the USD exchange rate.
+9% to +15% (non-USD)
What changed. Microsoft moved to a twice-yearly currency-alignment cadence and raised local-currency Azure pricing to track USD: GBP up 9%, DKK/EUR/NOK up 11%, SEK up 15%.
Who it hits. Any pay-as-you-go or CSP customer billed in GBP, EUR, DKK, NOK, or SEK. Enterprise Agreement customers are shielded up to their contract's price-protection window.
What to do
Check your billing currency and agreement type, and rebuild your forecast around the new rate instead of the one you budgeted with.
A new entry appears when a provider changes a price we track. Read it in any feed reader. Follow the whole log, one topic, or one provider. Want the findings instead of the feed? See the price change index.
Every entry has one primary source: the provider's own blog post, What's New announcement, changelog, or pricing page. We fetched and read every one ourselves. If we couldn't find a provider-published figure with a real date attached, the entry doesn't ship, no matter how widely it was reported elsewhere.
Why include price cuts? Doesn't that undercut the point of a cost-monitoring company?
A log that only shows increases is marketing with a straight face. Providers cut prices about as often as they raise them, and a cut is still something worth knowing about. It changes whether a workload you moved off a provider to save money is worth moving back.
How is this different from a provider's own What's New feed?
AWS, Azure, and the rest each publish their own firehose, and none of them cross-reference the others. This pulls only the entries that moved a bill, across every provider CostMon monitors, with the same fields every time: what changed, who it hits, and what to do about it.
Can I get notified when a new entry is added?
Subscribe to the RSS feed at /price-changes.xml. Every new entry is a dated item the moment it ships, so you don't have to keep checking back.
A provider changed something you haven't logged yet. Can I tell you?
Yes. Send the announcement and its source link through the contact page. If it's a real, provider-sourced price change, it gets added with the same sourcing standard as everything else here.