The Price Sheet You Priced Your Product On Is Already Stale
If you launched an AI feature in the spring and set your own pricing from the vendor's rate card, there is a good chance that rate card has changed twice since. One third-party tracker that follows text-model API prices counted at least 28 price changes across Anthropic, OpenAI, Google, DeepSeek and xAI between July and September. Most new flagship models arrived at or below the price of the model they replaced. One vendor raised prices. And several prices now come with an end date.
I am citing that tracker as a secondary source, so treat its exact figures as a snapshot and check each vendor's pricing page before you act on a number. The pattern is what matters, and the pattern is consistent across every source I looked at: prices move often, in both directions, and promotional rates expire.
This is mostly good news. Cheaper models mean wider margins for products already in market. But a founder who treats the vendor's price as a constant is exposed on both sides. If prices fall, a competitor with a flexible architecture passes the savings to customers and undercuts you. If a promotion ends, your gross margin drops overnight and your pricing page does not.
This article is about the second problem. Our posts on LLM API pricing compared and managing LLM API costs cover the rate cards and the standard tactics. Here the question is structural: how do you design a product so that a price change at your vendor is a config edit rather than an incident.
Four Kinds of Price Change, and Which Ones Hurt
Not every change is the same. Lumping them together leads to the wrong fix. In the last quarter we saw four distinct patterns, and each needs a different response.
New Model, Same or Lower Price
The most common pattern. A vendor ships a newer model at the price of the old one, or cheaper. For example, per the tracker, OpenAI launched a GPT-6 Sol model in September priced at $2 per million input tokens and $10 per million output tokens, well below the $5 and $30 standard rate for GPT-5.6 Sol it had launched in July. This is a benefit you can only capture if you can switch models without a rewrite.
Price Cuts on Existing Models
Vendors also cut prices on models already in market. In late July, per the same tracker, OpenAI cut its small GPT-5.6 Luna model by 80%. If your costs are indexed to the old price, you are overpaying until you notice.
Promotional Prices With an End Date
This is the one that hurts. As of early October, the tracker lists a promotional rate on GPT-5.6 Sol that runs at least until November 21, and introductory pricing on Google's Gemini 3.6, 3.7 and 3.8 Flash models through December 31, rising to double the rate on January 1, 2027. If your margin model was built on the promotional number, you have a date on the calendar when it breaks.
Changes in Billing Structure
The least discussed pattern. DeepSeek moved in August from a flat rate to peak and off-peak billing, with off-peak at half the peak rate. A change like that does not show up as a headline price increase, but it changes what your bill looks like depending on when your users are active. A product whose traffic peaks in business hours in one time zone pays differently than a batch workload that can run at night.
The lesson across all four: you need a way to see the price you are paying per request, a way to change the model behind a feature, and a way to schedule work. Nothing else about this problem matters if you cannot do those three things.
Put a Thin Abstraction Between Your Product and the Model
The first structural fix is boring and effective: do not call vendor SDKs from your feature code. Route every model call through one internal function or service that takes a task name and returns a result.
In practice this means your product code says "summarize this document" or "classify this support ticket," and a single layer decides which model handles it, what the prompt template is, what the timeout is, and what to do on failure. The feature code does not know whether it is talking to Anthropic, OpenAI or Google.
You do not need to build this from scratch. Open source and hosted gateways such as LiteLLM, OpenRouter, Portkey and Cloudflare AI Gateway give you a unified interface to many providers, along with logging, caching and rate limits. Which one fits depends on whether you want to self-host, how much logging you need and what your compliance posture is. We went through the trade-offs in our post on AI model routing and LLM cost optimization.
What the Layer Should Own
- A model map per task. A config file or table that says which model serves which task, so changing it needs no deploy of feature code.
- Prompt versions. Prompts often need small adjustments between models. Store them next to the model map.
- Fallbacks. If a provider is down or rate limited, the layer retries on a second model. This also protects you from outages, which is a reason to build the layer even if prices never moved.
- Cost logging. Every call records tokens in, tokens out, cached tokens, model and task, so you can compute real cost per feature.
The honest caveat: abstraction has a cost. Models differ in tool-calling behavior, output formatting and how they follow long instructions, so swapping one for another is never purely a config change. The layer reduces the work, it does not remove it. Budget a day or two of evaluation each time you move a production task to a new model.
Measure Cost Per Task, Not Cost Per Month
A monthly invoice from your AI vendor tells you nothing you can act on. A cost per task, tracked over time, tells you everything.
Define a unit that matches how you sell. For a support copilot, that might be cost per resolved ticket. For a document tool, cost per document processed. For a coding assistant, cost per accepted suggestion. Then log the model spend against that unit in your analytics.
With that number in hand, price changes become easy to read. If cost per resolved ticket falls from 12 cents to 4 cents after a model swap, you can see the margin gain. If a promotional rate ends and the number triples on a certain date, you saw it coming, because you compared the date to your forecast.
Build the Dashboard Once
A basic version takes a few days. Your gateway or logging layer writes each call to a table with feature, model, token counts, cached tokens and a computed dollar cost from a price table you maintain. A Metabase or Grafana board shows cost per task by day, by feature, and by customer. Add one alert: notify the team if cost per task moves more than 25% week over week.
The price table is the part founders skip. Keep it in your database with an effective-date column, so each row says "this model costs this much from this date until this date." When a vendor announces a change, you add a row. Historical costs stay correct, and forecasts can use future-dated rows for scheduled increases.
Include the long-context surcharges too. Some vendors price requests above a certain prompt length at a higher rate. According to the tracker, one OpenAI flagship charges more once a prompt passes 272K tokens. If your retrieval pipeline stuffs large contexts, that threshold can quietly double a request's cost.
Cut the Bill Structurally With Caching and Batching
Some savings do not depend on what the vendor charges. They come from sending less work to the model in the first place, and they hold up through price changes in either direction.
Prompt Caching
Most major vendors now discount cached input tokens heavily. The tracker lists cache hits at a small fraction of the normal input price on several recent models, including Anthropic's latest releases and OpenAI's GPT-6 line. If your prompts begin with a long, stable block, such as a system prompt, policy documents or a codebase summary, structure them so the stable part comes first and the variable part comes last. That ordering is what makes caching work. We wrote about this in more detail in prompt caching and smart routing.
Batch and Off-Peak Work
Anything a user is not waiting on should run in batch mode or off-peak. Nightly summaries, embeddings for new content, report generation and evaluation runs are all candidates. Most vendors offer a batch tier at a substantial discount, and with peak and off-peak billing appearing at providers like DeepSeek, scheduling is now a pricing lever as well as a performance one.
Route by Difficulty
Do not send every request to your best model. A cheap small model can classify intent, extract fields and handle simple questions. Escalate to the flagship only when a confidence check fails or the task needs deeper reasoning. Teams that do this well often move 60 to 80% of requests to the cheap tier, in our experience, with a small quality loss that a good evaluation set can measure.
Trim What You Send
Retrieval that returns twelve chunks when four would do, chat histories that resend the whole conversation, and verbose JSON schemas all inflate input tokens. A pass over your highest-volume prompts for dead weight is often the quickest win, and it pays off under every vendor's price list.
Price Your Own Product With a Buffer
The other half of the problem is on your side of the table. Your customers pay you a price, and your vendor charges you a cost, and the gap between them is your margin. If the cost can change by a factor of two on a known date, your pricing needs slack.
Three rules we give founders:
- Model margin at the post-promotion price. If a promotion ends on a date, assume the standard rate in your plan. A promotional discount is upside, not a base case. This one rule would have saved several teams we have talked to from a bad month.
- Keep a gross margin floor for the AI line. Many AI products aim for 60 to 70% gross margin on the AI portion of cost of goods. Pick your floor and alert on it. If the model cost per task breaches it, the team swaps a model or adjusts a limit before the next billing cycle.
- Avoid fixed promises tied to a model. Do not advertise "unlimited GPT-whatever" in your plan. Sell outcomes and fair-use limits, so you can change the engine underneath without breaking a promise.
Usage-based pricing deserves a mention. If your price scales with usage, your revenue rises with your costs and a vendor increase is partly absorbed. If you sell flat subscriptions with heavy users, you carry the risk. Credits, fair-use caps and overage tiers are the standard tools. Pick the one that matches how your customers already think about value.
A Short Checklist for This Week
You can do most of this in a week with one engineer. Here is the order I would use.
- Open each vendor's pricing page and list every model you call, its current price and any announced end date. Put the dates in a shared calendar.
- Add a price table with effective dates and compute dollar cost on every logged call.
- Route all model calls through one internal layer, even if it only wraps one provider today.
- Build or buy an evaluation set of 50 to 200 real examples per task, so you can test a cheaper model in an afternoon instead of arguing about it.
- Reorder your prompts so stable content comes first and turn on caching.
- Move non-urgent work to batch or off-peak.
- Re-run your margin model at post-promotion prices and fix any plan that fails.
None of this is glamorous, and all of it compounds. The founders who handle the next round of price changes calmly will be the ones who spent a week on plumbing before it happened.
If your AI feature is live and you are not sure what each request really costs, or you want a team to build the routing layer and cost dashboard for you, book a free strategy call and we will look at your numbers together.