DeepSeek pushed the official version of its V4-Flash model into public beta on Friday, and independent testing quickly confirmed something unusual: the Chinese lab's small, cheap model now outscores its own flagship. Benchmarking firm Artificial Analysis gave the update, labeled DeepSeek-V4-Flash-0731, a score of 50 on its Intelligence Index, a 10-point jump over the version released in April 2026 and 6 points ahead of the larger, far more expensive DeepSeek V4 Pro.
The timing sharpens the stakes. The release landed one day after OpenAI cut the price of GPT-5.6 Luna, its budget frontier model, by 80% under pressure from cost-conscious customers. Even after that cut, Artificial Analysis estimates the new DeepSeek model sits just 1 point behind Luna in intelligence while costing roughly 60% less per task on DeepSeek's own API.
The budget tier catches the flagship
When DeepSeek launched its V4 lineup in April, it followed the industry's standard two-tier logic: Pro, a 1.6T-parameter model, for maximum capability, and Flash for speed and low cost. Three months later, that hierarchy has inverted, at least temporarily. The V4 Pro API and the models behind DeepSeek's consumer app have not changed; only the Flash API got the upgrade.
The company's own framing of what changed is direct. Its update log states the release has "significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview." Agent tasks are jobs a model completes on its own across many steps, operating a terminal, editing code, and calling other software, rather than answering a single question. They are the category the industry currently believes the money is in, which is why a budget model beating a flagship preview at them is more than a leaderboard curiosity.
Just as striking is what did not change. The model keeps the exact architecture and size of the April version: 284B total parameters, the internal settings that store what a model knows, with only 13B active for any given word, which keeps running costs low. DeepSeek says the gains came entirely from post-training, the refinement stage applied after a model's core knowledge is built. Pricing is untouched at $0.14 per 1M input tokens and $0.28 per 1M output tokens, and DeepSeek discounts cached input, text the system has already processed, by 98%, versus the 90% most rivals offer.
Independent numbers support the agent claim. On GDPval-AA v2, Artificial Analysis's test of real-world work tasks, the model's rating rose to 1559 from 1189, and Terminal-Bench 2.1, which measures how well a model operates a command line, climbed 17 points to 79% in the firm's testing.
What the score does not settle
The model still invents answers at a high rate. On Artificial Analysis's knowledge test, when it does not know something, it fabricates a response 84% of the time. That is 12 points better than its predecessor, but the entire improvement came from declining to guess more often; the share of questions it actually gets right did not move.
It is also verbose. The model produced roughly 206M output tokens to complete the Intelligence Index, 12% fewer than the April version but still around double the median for comparable models. In long agent loops, that means real bills run meaningfully higher than the rock-bottom per-token price suggests.
And the ceiling has not been reached. Anthropic's Claude Opus 4.8 still leads on every benchmark DeepSeek itself published, and among open models Kimi K3 remains 7 points ahead on the index. The model handles text only, no images, and its weights, the files that would let anyone run it themselves, are expected but not yet released, so for now it exists only as a paid API.
DeepSeek says the official version of V4-Pro "will follow soon." If a post-training pass alone lifted the small model 10 points, the open question is what the same treatment does to a model nearly 6 times its size. The price of near-frontier intelligence fell twice this week. Only one of those drops required touching a price list.
