OpenAI and Anthropic are slashing prices on their mid-tier AI models as cheaper Chinese alternatives gain traction among cost-conscious businesses and developers. The price cuts mark a shift from competing purely on model performance to competing on cost, with Chinese labs like DeepSeek and Moonshot forcing the market down.

OpenAI cut prices on GPT-5.6 Luna, its fastest and most affordable model, by 80 percent. Input tokens dropped from $1 to $0.20 per million tokens, while output tokens fell from $6 to $1.20 per million. Anthropic responded by launching Claude Opus 5 at half the price of its most capable model, Fable 5, at $5 per million input tokens and $25 per million output tokens. It also cancelled a planned price increase for the Sonnet 5 model scheduled for September.

The moves come as companies like DoorDash and Airbnb have reportedly started using Chinese-made models to reduce their AI operating costs. Moonshot’s Kimi K3 and DeepSeek’s V4 Flash now offer competitive performance at significantly lower price points, narrowing the gap that US labs once took for granted.

The numbers behind the cuts

Performance benchmarks from Artificial Analysis show how close the competition has become. Anthropic’s Opus 5 running at medium effort roughly matches Moonshot’s Kimi K3 at max effort in both performance and cost per task. OpenAI’s GPT-5.6 Luna at max effort performs similarly to DeepSeek’s V4 Flash at max effort, but Luna costs nearly twice as much per task even after the cuts.

This pricing pressure is structural, not temporary. Chinese models are largely open-weight, meaning developers can download and run them on their own hardware without ongoing API fees. That makes them attractive for high-volume use cases where the cost difference adds up quickly. A business processing 100 million tokens a day saves serious money switching from Luna at $0.20 input to DeepSeek V4 Flash at roughly half that, especially when running on their own GPU infrastructure.

What this means for developers

If you are building on any of these APIs, the price war is good news. The immediate effect is cheaper inference, which makes AI features more viable in cost-sensitive applications like customer support triage, content generation pipelines, and background data processing. The medium-term effect is more competition on capabilities rather than just cost — labs that cannot differentiate on performance will have to keep cutting prices or get squeezed out.

Mantas Lukauskas, AI tech lead at Hostinger, described the dynamic as: “The US labs have cut the middle and are defending the top.” The top-tier frontier models still command premium pricing for complex reasoning tasks where quality matters more than cost, but the mid-range is rapidly becoming a commodity market.

One complication is that headline token prices do not tell the whole story. More capable models may need fewer tokens or fewer attempts to complete a task, which can make a higher per-token price cheaper overall. Models also operate at different effort settings that affect both output quality and compute cost. Comparing price per task rather than price per token gives a more accurate picture, but that requires benchmarking against your own workload rather than relying on published rates.

Broader industry pressure

The price cuts come alongside a broader shift in how US labs charge for their services. Both OpenAI and Anthropic are moving enterprise customers from flat subscriptions to usage-based billing, which increases costs for high-volume users even as per-token rates fall. The combination of lower headline prices but more granular billing creates a confusing procurement environment for organizations trying to forecast AI spending.

Investor pressure also plays a role. With both companies reportedly planning IPOs at trillion-dollar valuations, they need to show that massive AI infrastructure spending can translate into sustainable revenue growth. Price cuts that drive adoption help build that revenue base, even if margins on individual transactions shrink.

The open-source factor

What makes this price war different from previous rounds of cloud pricing competition is the open-weight dynamic. DeepSeek and Moonshot release model weights that anyone can download, self-host, and fine-tune. A company running high-volume inference can avoid API fees entirely by running a Chinese model on its own hardware. That puts a hard ceiling on how much US labs can charge for mid-tier models — if the price goes too high, customers leave for self-hosted alternatives.

This is not purely about cost. Self-hosting gives organizations full control over data privacy, latency, and availability. For regulated industries like finance and healthcare where sending data to a third-party API is problematic, open-weight models are increasingly the default choice. The US labs are betting that convenience, reliability, and ecosystem integration will keep customers on their APIs even when cheaper alternatives exist, but that bet gets harder to sustain as the performance gap narrows.

What comes next

The most likely near-term outcome is continued price compression across the mid-tier market, with US labs competing on developer experience, brand trust, and integration with existing cloud platforms rather than raw price. The frontier models at the top of each lab’s lineup will retain premium pricing for the workloads that genuinely need them — complex reasoning, long-context analysis, and specialised agentic tasks — but everything else becomes a volume business.

For developers, the practical advice is to benchmark against your actual workload rather than relying on published token prices. Run the same task across multiple providers at different effort levels and measure total cost per completed task. The answer will depend heavily on what you are building, and the cheapest provider on paper may not be the cheapest in practice once throughput, latency, and reliability are factored in.

Leave a Reply

Your email address will not be published. Required fields are marked *