Use a Temporary AI Price Cut to Improve Margin
Treat a temporary AI price cut as a margin test, not a pricing reset, and bank savings before passing any discount to customers.

A temporary cut is a test, not a new list price
Your next AI invoice is smaller, and the tempting move is to pass the saving to customers. The safer move is to treat the window as a margin test, not a pricing reset.
OpenAI reduced the API price for GPT-5.6 Sol on August 21, 2026, from $5 per million input tokens and $30 per million output tokens to $4 per million input tokens and $20 per million output tokens. OpenAI committed the promotional GPT-5.6 Sol pricing to remain available at least through November 21, 2026. A standard GPT-5.6 Sol workload using 2 million input tokens and 500,000 output tokens would cost $25 under the pre-cut rates and $18 under the August 2026 rates, a 28% reduction. That gives you a short period to measure what the savings do before the promo ends.
Profit is a number; cash pays the invoice. A lower API bill improves both only if you keep the savings where they can cover the next bill, not just the current one.
Run the test before you touch customer prices
- Recompute per-customer AI cost at old and new rates. Pull a recent month of token usage by customer or plan, split input from output, and multiply by the old and new rates. The result is a small table: customer, average monthly input, average monthly output, old cost, new cost, and the delta.
- Add the edge cases before you celebrate. A GPT-5.6 Sol request whose input exceeds roughly 272,000 tokens bills at 2x the input rate and 1.5x the output rate for the entire request. A customer who pastes a whole codebase, a long contract, or a large log file can erase the discount in a call. End with a worst-case monthly cost for your top customers.
- Use the provider's example as a check. A real mix heavier on output may save more; a mostly input mix may save less than the headline suggests.
- Bank the savings in a named promo buffer. Hold enough to cover the month after expiry, your worst recent long-context month, and a small reserve for usage spikes. Keep the buffer visible on the bank balance, not buried in general operating cash. Classify each saving as banked margin or customer pass-through before you move any price.
If your product bundles AI with other features, allocate the API cost to the feature that triggers it. Do not let the whole product take a cost that only that workflow creates. A workflow that triggers long-context requests should carry the long-context surcharge, not the rest of the product.
The offer must survive the return to standard rates
Pass savings only when they buy durable usage. Make a time-limited offer that requires a higher tier, added seats, annual billing, or a usage commitment, and that buys a customer behavior you can measure. The offer is done when you can name the cohort: who accepted, what changed, and whether contribution margin improved after the offer.
Set the price floor before the offer: post-promo cost plus your target margin. Do not ship an offer that would not survive the return to standard rates. State the offer explicitly: available through a date, tied to the offer conditions, and reversible when those conditions end.
A list-price cut tells every customer that the old price was too high, even when only your input cost changed. Slower growth is a deliberate choice when the alternative is a discount you cannot afford.
Watch the cohort for a full billing cycle after the offer. Keep the offer terms, usage, and margin in one note. Check whether the customers kept the commitment and used the feature more often. Improved margin is evidence for the next offer; flat margin is a reason to keep the price where it was.