<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>CostMon — AI Price Changes</title><description>Every AI price change CostMon tracks that actually moved a bill, cuts and increases both, with the source and what to do about it.</description><link>https://costmon.com</link><item><title>OpenAI cut GPT-5.6 Luna&apos;s price 80% and Terra&apos;s 20%, leaving the flagship tier untouched.</title><link>https://costmon.com/price-changes#gpt-5-6-luna-terra-price-cut</link><guid isPermaLink="true">https://costmon.com/price-changes#gpt-5-6-luna-terra-price-cut</guid><description>Citing inference-stack efficiency gains, OpenAI dropped Terra to $2/$12 per million tokens and Luna to $0.20/$1.20, while the top-tier Sol model&apos;s price stayed flat. (Luna −80%, Terra −20%) — Confirm the lower rate appears on your next invoice, then check whether tasks still pinned to Sol can safely move down to the now much cheaper Luna or Terra.</description><pubDate>Thu, 30 Jul 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>Cut</category></item><item><title>GPT-5.5 launched priced higher than GPT-5.4, betting token efficiency would offset it.</title><link>https://costmon.com/price-changes#gpt-5-5-price-increase</link><guid isPermaLink="true">https://costmon.com/price-changes#gpt-5-5-price-increase</guid><description>GPT-5.5 launched at roughly double GPT-5.4&apos;s per-token rate. OpenAI&apos;s own announcement frames this as a deliberate increase, offset by claimed gains in token efficiency per task. ($5 / $30 per M tokens) — Run a side-by-side cost comparison on real workloads before moving production traffic. Confirm the token savings offset the higher headline price for you.</description><pubDate>Fri, 24 Apr 2026 00:00:00 GMT</pubDate><category>OpenAI</category><category>Increase</category></item><item><title>Gemini 3 Flash turned on context caching by default, cutting repeated-token cost up to 90%.</title><link>https://costmon.com/price-changes#gemini-3-flash-caching-batch-pricing-2025</link><guid isPermaLink="true">https://costmon.com/price-changes#gemini-3-flash-caching-batch-pricing-2025</guid><description>Gemini 3 Flash launched at $0.50 per million input tokens and $3 per million output tokens, with context caching standard rather than opt-in, plus a 50% discount for asynchronous Batch API jobs. (up to −90% w/ caching) — Confirm caching is active on your repeated prompts, and route non-interactive jobs through the Batch API to capture the 50% discount.</description><pubDate>Wed, 17 Dec 2025 00:00:00 GMT</pubDate><category>Google (Gemini)</category><category>Cut</category></item><item><title>Claude Opus 4.5 cut Opus-tier pricing to $5/$25 per million tokens.</title><link>https://costmon.com/price-changes#claude-opus-4-5-pricing-cut</link><guid isPermaLink="true">https://costmon.com/price-changes#claude-opus-4-5-pricing-cut</guid><description>Opus 4.5 launched at $5 input / $25 output per million tokens, a cut Anthropic frames as putting frontier-tier capability within reach of more teams. ($5 / $25 per M tokens) — Re-evaluate workloads kept on Sonnet purely for cost. At the new rate, Opus may be affordable for tasks where the extra capability wasn&apos;t worth the old premium.</description><pubDate>Mon, 24 Nov 2025 00:00:00 GMT</pubDate><category>Anthropic</category><category>Cut</category></item><item><title>Gemini 3 Pro launched as Google&apos;s new flagship, at a higher output rate than 2.5 Pro.</title><link>https://costmon.com/price-changes#gemini-3-pro-launch-pricing-2025</link><guid isPermaLink="true">https://costmon.com/price-changes#gemini-3-pro-launch-pricing-2025</guid><description>Google&apos;s new flagship reasoning and agentic model launched at $2 per million input tokens and $12 per million output tokens for prompts up to 200K tokens. ($2 / $12 per M tokens) — Benchmark real task cost against your current model before switching. The higher output rate needs to be earned back in fewer tokens per task, not assumed.</description><pubDate>Tue, 18 Nov 2025 00:00:00 GMT</pubDate><category>Google (Gemini)</category><category>Restructure</category></item><item><title>Claude Haiku 4.5 launched at a third of Sonnet 4&apos;s price for comparable coding performance.</title><link>https://costmon.com/price-changes#claude-haiku-4-5-pricing</link><guid isPermaLink="true">https://costmon.com/price-changes#claude-haiku-4-5-pricing</guid><description>The new small model launched at a flat $1 input / $5 output per million tokens, positioned by Anthropic as Sonnet-4-level coding at a third of the cost. ($1 / $5 per M tokens) — Benchmark current Sonnet-4-class workloads against Haiku 4.5 and shift eligible traffic down a tier, checking quality on your own tasks first.</description><pubDate>Wed, 15 Oct 2025 00:00:00 GMT</pubDate><category>Anthropic</category><category>Restructure</category></item><item><title>GPT-5 replaced the GPT-4o/o-series lineup with one model family sold at three price points.</title><link>https://costmon.com/price-changes#gpt-5-launch-pricing</link><guid isPermaLink="true">https://costmon.com/price-changes#gpt-5-launch-pricing</guid><description>GPT-5 launched at $1.25 per million input tokens and $10 per million output tokens, with mini and nano tiers underneath it: one family instead of separate reasoning and non-reasoning models. ($1.25 / $10 per M tokens) — Map existing GPT-4o and o3 call sites to gpt-5, gpt-5-mini, or gpt-5-nano by task difficulty, and re-benchmark spend rather than assuming a like-for-like swap.</description><pubDate>Thu, 07 Aug 2025 00:00:00 GMT</pubDate><category>OpenAI</category><category>Restructure</category></item><item><title>Gemini 2.5 Flash&apos;s stable release raised input price but cut output price and merged two rate tiers into one.</title><link>https://costmon.com/price-changes#gemini-2-5-flash-stable-repricing-2025</link><guid isPermaLink="true">https://costmon.com/price-changes#gemini-2-5-flash-stable-repricing-2025</guid><description>Going stable, Google removed 2.5 Flash&apos;s separate thinking vs. non-thinking price tiers, raising input from $0.15 to $0.30 per million tokens while cutting output from $3.50 to $2.50. (in +100%, out −29%) — Recompute cost-per-request at the new $0.30/$2.50 blended rate before migrating off the preview model. Input-heavy workloads got pricier even though output-heavy ones got cheaper.</description><pubDate>Tue, 17 Jun 2025 00:00:00 GMT</pubDate><category>Google (Gemini)</category><category>Restructure</category></item><item><title>Anthropic cut tool-use output tokens on Claude 3.7 Sonnet by up to 70%.</title><link>https://costmon.com/price-changes#anthropic-token-saving-updates-2025</link><guid isPermaLink="true">https://costmon.com/price-changes#anthropic-token-saving-updates-2025</guid><description>A new token-efficient tool-use mode cut output tokens consumed on tool-calling workloads, and prompt-cache reads stopped counting against input-token-per-minute rate limits. (up to −70% output tokens) — Turn on token-efficient tool use for tool-calling workloads, and check whether old throttling logic built around the rate limit is still needed.</description><pubDate>Thu, 13 Mar 2025 00:00:00 GMT</pubDate><category>Anthropic</category><category>Cut</category></item><item><title>Anthropic launched a batch endpoint priced at half the standard API rate.</title><link>https://costmon.com/price-changes#anthropic-message-batches-api-50-off</link><guid isPermaLink="true">https://costmon.com/price-changes#anthropic-message-batches-api-50-off</guid><description>The Message Batches API processes up to 10,000 queries asynchronously within 24 hours, at half the price of a standard synchronous call. (−50% vs standard API) — Audit API traffic for anything without a real-time latency requirement and move it to the Batches API to cut that portion of the bill in half.</description><pubDate>Tue, 08 Oct 2024 00:00:00 GMT</pubDate><category>Anthropic</category><category>Restructure</category></item><item><title>GPT-4o launched at half the price of GPT-4 Turbo, with 5x the rate limit.</title><link>https://costmon.com/price-changes#gpt-4o-half-price-vs-gpt-4-turbo</link><guid isPermaLink="true">https://costmon.com/price-changes#gpt-4o-half-price-vs-gpt-4-turbo</guid><description>OpenAI&apos;s new flagship model replaced GPT-4 Turbo as the default, matching or beating its quality at half the per-token price. (−50% vs GPT-4 Turbo) — Re-point integrations at gpt-4o and re-run your eval suite before rolling it out. A cheaper model at the same price tier can still shift output style.</description><pubDate>Mon, 13 May 2024 00:00:00 GMT</pubDate><category>OpenAI</category><category>Cut</category></item></channel></rss>