DeepSeek is raising API prices for its V4 model family, with increases ranging from roughly 50% to more than 1,100% depending on model, tier, and time of day. The new pricing takes effect August 16 and introduces peak and off-peak rates, with off-peak usage priced at half the peak rate.
The increase lands hardest on developers who run workloads during peak hours. But the company's cache-hit discounts and time-of-day pricing mean the effective cost increase depends heavily on how and when applications call the API.
Flash and Pro pricing breakdown
DeepSeek V4-Flash now costs $0.22 per million input tokens (cache miss) and $0.66 per million output tokens off-peak. At peak, those prices rise to $0.44 and $1.32. That represents a 57% to 214% increase on inputs and a 136% to 371% increase on outputs versus the previous flat rate.
V4-Pro is priced at $0.66 per million input tokens (cache miss) and $1.98 per million output tokens off-peak, rising to $1.32 and $3.96 at peak. That's a 51% to 203% increase on inputs and 127% to 355% on outputs. Inputs with cache hits - when apps reuse stored prompts - see the steepest increases, ranging from 52% to 1,100%.
The pricing change was announced alongside the general availability of DeepSeek V4-Pro and the beta release of V4-Flash. Both models add flexible reasoning levels (low, high, max) and thinking modes that use chain-of-thought reasoning to improve answer accuracy.
Peak pricing changes the competitive picture
"On paper, at peak, against the right comparator, DeepSeek's price advantage does disappear, and in places inverts," said Sanchit Vir Gogia, chief analyst at Greyhound Research. "But in practice, the schedule's own clock and cache hand most of it back to any buyer paying attention."
At peak pricing, V4-Flash loses its cost edge over OpenAI's GPT-5.6 Luna, but the advantage holds off-peak, said Mark Tauschek, VP of research fellowships at Info-Tech Research Group. OpenAI has cut Luna's API pricing by 80% for off-peak use. DeepSeek V4-Pro still beats OpenAI's mid-tier Terra model on price even at peak, Tauschek said.
Gogia said V4-Flash's overall cost-per-task advantage over Luna shrinks from roughly sevenfold off-peak to 1.4 times at peak. The company's approximately 98% cache-hit discount - versus an industry norm nearer to 90% - has been a key mechanism keeping costs low. "The cache is where the advantage genuinely erodes," he said.
DeepSeek is steering usage to off-peak hours
DeepSeek said the pricing change is intended to "allocate resources more reasonably" and encourage customers to "schedule their tasks based on actual usage." In effect, 17 of every 24 hours stay at half price, which makes task timing an economic decision, Gogia said.
Western customers, who largely operate outside DeepSeek's home-market peak hours, likely absorb less of the increase. "The schedule re-prices exactly that mechanism," Gogia said.
The underlying driver is supply and demand, not margin expansion, Tauschek said. "This wasn't unexpected at all. When demand goes up, pricing goes up, because supply becomes constrained." Anthropic also raised API prices for the same reason in April. Third-party providers have not yet followed, but they will have to, he said.
For enterprises, the increase is unlikely to change adoption decisions, Tauschek said. Many US enterprises do not use DeepSeek at all due to policy constraints. Developers building on DeepSeek's API will feel the impact more directly, though the models remain cheaper than most alternatives.
Why this matters for IT and development teams
Developers building agentic workloads should treat DeepSeek's new pricing schedule as a routing variable, not just a cost line. The cache-hit discount and off-peak rates change the economics of when to call the API, and multi-model routing means these models can be swapped for alternatives without rewriting the application layer.
Gogia said the broader trend is toward substitutable model intelligence. Open weights, compatible APIs, and competitive pricing mean foundation model vendors no longer own the entire dependency stack. The lasting effect of DeepSeek, he said, is that "every provider must now explain why intelligence should command a premium once near-equivalent capability is available through several technical and commercial routes."
For teams using DeepSeek internally or in production, check your call patterns against peak hours and audit cache-hit rates - the model's cost per task can still beat competitors, but only if scheduling and caching are factored into deployment decisions. New to DeepSeek or model selection? Explore DeepSeek Courses has ROI-focused videos; and AI and Development has job-specific courses and plans drawn from developer workflows.
Your membership also unlocks: