DeepSeek has been the price floor of the AI industry since its R1 model landed in January 2025. Its rates are low enough that developers route high-volume work to it by default, and Chinese models now carry the majority of traffic on OpenRouter, the marketplace where developers shop for models across vendors. On Wednesday the Hangzhou company changed that bargain. It announced the general release of DeepSeek-V4-Pro, its flagship, and in the same post published a new API price list.

The new rates take effect at 16:00 UTC on August 16. Increases run from 50% to 1,100% above current prices depending on the model, the type of token and the hour of the day, a range Reuters calculated from the company's statement. Output on V4-Pro, meaning the text the model writes, goes from $0.87 per million tokens to $1.98 for most of the day and $3.96 during a seven-hour peak window.

None of this was a secret in outline. DeepSeek told developers on August 6 that a broad increase was coming and to plan accordingly, without saying how much or when. The answer arrived a week later, attached to a product launch.

The company frames the change as a discount. "Off-peak rates are 50% lower than peak," DeepSeek wrote in its launch announcement, calling it a way to schedule workloads more flexibly. The reference point there is the new peak rate, not the old flat one. Until now, a million output tokens on V4-Pro cost the same at any hour of any day. From Sunday they cost roughly 2.3 to 4.6 times more than they do today, and which end of that you pay depends on when your job runs.

Peak and off-peak pricing arrives August 16

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, per DeepSeek's API pricing page. Everything else is off-peak. The cheaper V4-Flash model moves too: input from $0.14 to $0.22 off-peak and $0.44 at peak, output from $0.28 to $0.66 and $1.32.

The steepest rise lands on cached input, the discounted rate charged when a request repeats text the system has already processed. On V4-Pro that goes from $0.003625 per million tokens to $0.022 off-peak and $0.044 at peak, a jump of more than 1,100% at the top end. Agent software, which resends the same instructions on every loop, is precisely the workload that leans on caching hardest.

What V4-Pro-0813 scores on DeepSeek's own tests

The model behind the endpoint is DeepSeek-V4-Pro-0813, ending a preview that began April 24. It uses a mixture-of-experts design with 1.6 trillion total parameters and 49 billion active for any given token, reads up to 1 million tokens of context and writes up to 384,000. It is live in the app under "Expert Mode," supports OpenAI's Responses API with one-click Codex setup, and offers reasoning effort settings from low to max. Its cheaper sibling had already outscored the flagship preview on agent tests in July.

DeepSeek's published table puts V4-Pro-0813 at 87.9 on Terminal Bench 2.1, which measures how well a model drives a command line, behind Kimi-K3 at 88.3 and Claude Fable 5 at 88.0. It leads two of the ten rows: 83.3 on Cybergym, a security benchmark, and 31.8 on the public AutomationBench. On Humanity's Last Exam it scores 42.7 without tools against Fable 5's 53.3. A footnote says the coding-agent numbers came from an unreleased in-house test harness and may vary under other frameworks. No outside lab has verified any of it.

The cost gap narrows but does not close

Even after Sunday, DeepSeek remains far below what it benchmarks against. Claude Fable 5 lists at $10 and $50 per million input and output tokens. xAI released Grok 4.6 the same day as V4-Pro at $2 and $6, among the cheapest Western frontier options. V4-Pro at full peak rates is $1.32 and $3.96.

What DeepSeek is testing is whether the developers who made it the most-used model family on the open market were buying capability or buying cheapness. If they stay, every other lab gets room to stop discounting. The floor just moved, and everyone above it will notice.