LLM API Price Changes & Rate Adjustments
Track official token rate cuts, off-peak pricing discounts, and API rate adjustments across major providers. Find cost-saving opportunities for your workloads.
GPT Image 2 cached input pricing clarification for Responses API
OpenAI clarified that cached input rates for GPT Image 2 and GPT Image 2.5 only apply to images generated via the Responses API. This is a documentation clarification rather than a pricing change, but it may affect cost expectations for users on other endpoints.
OpenAI adds FedRAMP 10% pricing uplift and Agents SDK/ChatKit docs
OpenAI's pricing page now notes FedRAMP endpoints cost 10% more than standard model rates, and adds Agents SDK and ChatKit references. FedRAMP users should expect higher bills; others see no pricing change.
Sora-2-pro video pricing table removed from OpenAI pricing page
The per-second pricing table for Sora-2 and Sora-2-pro video generation models has been removed from the OpenAI API pricing page, replaced with generic fine-tuning billing text. Pricing for these models is no longer documented here, which may indicate a restructure or removal of published rates.
Google AI audio pricing doubles starting January 1, 2027
Google published audio pricing tiers showing rates through Dec 31, 2026, then doubling on Jan 1, 2027. Multiple audio SKUs are affected, with per-10s rates rising 100% after the deadline.
Gemini 3.8 Flash TTS pricing introduced with 2027 rate increase
Google published pricing for the new gemini-3.8-flash-tts audio model: $0.50/$9.00 per unit through 2026, doubling to $1.00/$18.00 on Jan 1, 2027. Input caching also doubles. Teams should budget for the scheduled 100% increase.
OpenAI o4-mini pricing page adds WebRTC with WARP entry
The o4-mini pricing table now lists a 'WebRTC with WARP' option at $100.00/hour alongside the existing o4-mini-2025-04-16 rates ($4.00 input / $1.00 output). The prior data-sharing row and a $16.00 figure were removed, indicating a pricing/offering restructure.
OpenAI adds per-minute pricing for realtime translation and transcription models
OpenAI's pricing page now lists per-minute rates for gpt-realtime-translate ($0.034/min), gpt-live-transcribe and gpt-realtime-whisper ($0.017/min), plus gpt-transcribe ($0.0045/min). This introduces new usage-based pricing for realtime audio models.
OpenAI pricing page restructured; gpt-realtime-translate rates removed
The OpenAI API pricing page diff shows the gpt-realtime-translate live translation rate ($0.034/min) and related transcription model pricing rows removed, alongside a language selector and navigation update. This may indicate a page restructure or pricing table relocation rather than a confirmed pri
OpenAI o4-mini pricing page cleanup; token rates removed
The o4-mini pricing page dropped its per-token rates ($4.00 input / $1.00 cached / $16.00 output per 1M tokens) and now only shows a $100.00/hour figure, alongside removal of 'Meetups' links. This obscures standard token pricing for o4-mini users.
DeepSeek V4 Pro pricing page drops service continuation notice
The pricing page removed the note confirming DeepSeek V4 Pro API service would continue past September 14, 2026, and reworded off-peak hours to include weekends and Chinese public holidays. Concurrency limits (2500/500) are unchanged.
OpenAI o4-mini pricing page adds data-sharing tier and federation rules
The o4-mini pricing page now lists a 'with data sharing' variant alongside the standard o4-mini-2025-04-16 rate, and reorganizes workload identity federation content. This suggests a new discounted or alternate pricing tier tied to data sharing, which could affect cost and compliance decisions.
OpenAI o4-mini pricing page restructured; MCP section renamed
The o4-mini pricing table was reorganized: the 'MCP and Connectors' section is now 'MCP servers', and per-token input/cached/output prices ($4.00/$1.00/$16.00, and $2.00/$0.50/$8.00 with data sharing) were removed from the diff view. No confirmed price change for o4-mini itself.
OpenAI o4-mini pricing page adds Usage Insights entry
The o4-mini pricing page now lists a 'Usage Insights' line at $100.00/hour alongside the existing token pricing. This appears to be a new billing line item for usage analytics rather than a change to o4-mini token rates.
o4-mini pricing table adds cached input price column
The o4-mini-2025-04-16 pricing row now shows an additional $0.50 cached input rate alongside the existing $2.00 input and $100/hour figures. This is a minor documentation/table update rather than a change to headline pricing.
DeepSeek V4.1-Flash replaces V4-Flash; vision model removed
DeepSeek updated its pricing page: deepseek-v4-flash is now DeepSeek-V4.1-Flash with cache-hit input prices cut roughly 57% ($0.007 to $0.003 off-peak), and the deepseek-v4-flash-vision-exp model has been removed from the lineup.
Google AI pricing update: scheduled price increases effective Jan 1, 2027
Google AI pricing page now shows tiered pricing with current rates valid through Dec 31, 2026, and higher rates starting Jan 1, 2027. All listed prices double (e.g., $1.00 to $2.00, $5.00 to $10.00). This is a planned price increase for text, image, video, and audio modalities.
Google AI pricing page updated for Gemini, Veo, and Lyria models
The Google AI pricing page was updated on 2026-08-28, listing pricing for gemini-3.1-pro-preview, veo-3.1 preview models, and lyria-3 preview models. No specific price changes are detailed; only the page's last-updated timestamp changed.
OpenAI renames daybreak model aliases to gpt- prefix
OpenAI updated the pricing page for gpt-5, renaming the aliases 'daybreak-blue-latest' and 'daybreak-red-latest' to 'gpt-daybreak-blue-latest' and 'gpt-daybreak-red-latest'. These aliases still point to gpt-5.6-sol and gpt-5.6-sol (likely a typo for the second model). The change is cosmetic in docum
OpenAI gpt-5 pricing aliases updated for Daybreak program
OpenAI updated the gpt-5 pricing documentation to clarify that aliases like gpt-5.6-cyber will be updated to point to the latest models as new ones are released through the Daybreak program, with pricing adjusted accordingly. The change is minor, replacing 'frontier models' with 'models' and adjusti
Gemini 3.7 Flash pricing update with temporary discounts
Google introduces Gemini 3.7 Flash with promotional pricing through Dec 31, 2026, then doubling on Jan 1, 2027. Input, output, and storage prices are halved during the promo period. Existing pricing entries removed, indicating a new model tier.
Gemini Robotics-ER 1.6 Preview pricing restructured to per-request model
Google AI updated the pricing page for Gemini Robotics-ER 1.6 Preview, replacing the old per-prompt pricing with a per-request model. The free tier is now 5,000 search requests per month (shared across all Gemini 3.x models), with overage at $14 per 1,000 requests. The change applies to spatial reas
Google AI pricing page restructured with new plan details
Google AI pricing page has been updated with a new 'Google AI Plans' section, likely introducing new pricing tiers or plan options. This may affect cost calculations for users.
Google AI pricing update
Pricing changes detected on Google AI pricing page. May affect costs for API usage. Review new rates to manage budget.
Anthropic pricing page updated
Anthropic's pricing page has been updated, likely reflecting changes to model costs or rate structures. Users should review new prices to assess budget impact.
DeepSeek pricing update
Pricing page changed, likely reflecting new rates or model costs. Users should review for cost impact.
Gemini 3.1 Flash Lite Image pricing page updated
The pricing page for Gemini 3.1 Flash Lite Image was updated on 2026-07-09, with the model name changed to include '(Nano Banana 2 Lite) 🍌'. No pricing numbers changed in this diff, so this is a minor metadata update.
Anthropic pricing page updated
Anthropic's pricing page has changed, likely reflecting new model costs or rate adjustments. Users should review the updated pricing to understand any impact on their usage budgets.
Gemini 2.5 Flash Image pricing renamed to Nano Banana
Google updated the Gemini 2.5 Flash Image pricing line item, changing its display name from 'Gemini 2.5 Flash Image 🍌' to 'Gemini 2.5 Flash Image (Nano Banana) 🍌'. This appears to be a cosmetic rename with no pricing change, but may cause confusion for users matching SKUs.
Google AI Pricing Update for Live Translation
Google AI updated pricing for Gemini 3.5 Live Translate, increasing rates from $3.50 or $0.0053/min to $21.00 or $0.0315/min for audio input, with billing based on total input and output audio token consumption.
Google AI Pricing Update
The Google AI Pricing page has been updated to reflect the same information as the Standard page, with no changes to pricing. The update also includes a new feedback section and was last updated on 2026-06-02 UTC.
Gemini 3 Pro Pricing Update
Google updated the pricing information for Gemini 3 Pro Image on May 29, 2026. The update includes a new version, Gemini 3 Pro, replacing Gemini 3.1 Pro, with an updated feedback mechanism and the same date.
Gemini 3.1 Pro Image Pricing Update
Google AI added Gemini 3.1 Pro Image to its pricing page, replacing Gemini 3 Pro Image Preview. This change may impact users who were using the preview version.
Gemini 3.1 Flash-Lite Pricing Update
The pricing page for Gemini 3.1 Flash-Lite was updated on 2026-05-27 UTC, with a change in the feedback section and last updated date.
Claude Opus 4 Pricing and Model Updates
Anthropic updated pricing for Claude Opus 4.7, Claude Sonnet 4.6, and Claude Haiku 4.5 models. Pricing changes include $5/input MTok and $25/output MTok for Claude Opus 4.7, $3/input MTok and $15/output MTok for Claude Sonnet 4.6, and $1/input MTok for Claude Haiku 4.5. No output pricing is listed f
Claude Opus 4 Pricing Update
Anthropic updated the pricing for Claude Opus 4, changing input and output token costs. This may impact users with high usage or specific budget constraints.
Anthropic Updates Claude Opus 4 Pricing
Anthropic updated pricing and model information for Claude Opus 4 variants, including new model IDs and pricing details, particularly for Claude Haiku 4.5.
Google AI Pricing Update for Gemini 3.5 Flash
Google AI pricing page updated with new pricing for Gemini 3.5 Flash, including $0.08 and $0.27 per minute rates, and introduction of managed agents and environments.
Google AI Pricing Deploying App
Deploying your app section added to Google AI Pricing page, likely providing new information on deployment costs or considerations.
Claude Opus 4.7 Pricing and Model Details Update
Anthropic has updated the pricing and details for Claude Opus 4.7, Claude Sonnet 4.6, and Claude Haiku 4.5 models. The changes include new pricing tiers and model descriptions, with a focus on improved capabilities, especially in agentic coding for Claude Opus 4.7.
Google AI Pricing Update: Gemini API Change
Google AI pricing page updated, replacing VertexAI Gemini API with Gemini Enterprise Agent Platform Gemini API, potentially affecting access and costs.
Gemini 3.1 Flash-Lite pricing page update
The last updated timestamp on the Gemini 3.1 Flash-Lite pricing page has changed from 2026-05-06 UTC to 2026-05-07 UTC, indicating a minor update to the page.
Claude Opus 4 Pricing and Model Info Update
Anthropic updated Claude Opus 4 pricing and model information, introducing new models and changing existing ones' API IDs and pricing.
Google AI Interactions API Migration
Google AI is requiring migration to Interactions API, with breaking changes planned for May 2026, which may require code changes and impact project timelines.
Google AI Pricing Webhooks Update
Google AI added information on webhooks to its pricing page, potentially indicating new features or changes to notification mechanisms for AI usage.
Claude Sonnet 4 Pricing Update
Anthropic updated the pricing information for Claude Sonnet 4. The changes reflect modifications in how pricing details are presented, potentially indicating a change in cost structure or billing.
Claude Opus 4 Pricing and Model Details Update
Anthropic updated the pricing and details for Claude Opus 4 and other models. The max output for Claude Opus 4 changed from 128k tokens to 64k tokens. This change may impact users who rely on the higher output for complex tasks.
Anthropic Models Docs pricing and info update
Anthropic updated the max output and reliable knowledge cutoff information for their models. The max output for some models changed from 64k tokens to 128k tokens. The reliable knowledge cutoff dates also changed.
Max output updated in Anthropic Models Docs
The maximum output for Anthropic models has been updated from 100k to 128k tokens, which may impact users who rely on specific output limits for their applications.
Google AI Pricing Update: Gemini API Renamed
Google AI pricing page updated, renaming VertexAI Gemini API to Gemini Enterprise Agent Platform Gemini API, and Gemini Embedding 2 Preview to Gemini Embedding 2. Pricing information and links to specific pricing pages have been updated.
Anthropic Pricing Update for Claude Code
Anthropic updated its pricing page for Claude Code, which may reflect changes in cost or usage limits for the AI model.
Want to compare live rates across all vendors?
Open the Price Matrix to filter input, output, and cached token prices side by side.