AI model benchmarking firm Artificial Analysis published its score for DeepSeek’s latest iteration of the lightweight V4-Flash model this week, after it went into public beta on Friday. Surprisingly, the Chinese developer’s cheapest model scored 50 on the platform’s Intelligence Index, placing it 6 points above its flagship DeepSeek V4 Pro-Preview model.
DeepSeek also remains significantly cheaper than leading U.S. models. Competitors in the U.S. have taken steps to try to close this gap in recent months, most notably OpenAI, which slashed the cost of its own budget model by 80% just days after release. Regardless, at $0.14 per million input tokens and $0.28 per million output tokens, DeepSeek’s model is still 60% cheaper than GPT-Luna.
How Flash overtook Pro-Preview in intelligence
The update logs published by DeepSeek show that the latest version of Flash has the same fundamental architecture and size as its previous iteration from April this year.
The variable that changed was the post-training refinement. While the flagship V4-Pro-Preview API remains the same, the Flash API has received a major overhaul. The result is "significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview."
Agent capability is becoming an increasingly important benchmark for AI models, reflecting their expansion far beyond basic chatbot capabilities. This refers to a model’s ability to tackle complex, multi-stage tasks such as writing and testing code.
AI benchmarking firm Artificial Analysis corroborated DeepSeek’s findings regarding its latest model. In independent testing, Flash climbed 17 points over its predecessor on Terminal-Bench 2.1, a test that measures how well a model operates a command line.
On GDPval-AA v2 - Artificial Analysis's test of real-world work tasks across sectors like finance, engineering and healthcare - the model's rating rose to 1,559 from 1,189.
These findings suggest that V4-Flash’s agentic capabilities truly are a major step up.
Where DeepSeek V4-Flash still falls short
Although DeepSeek’s model has made marked improvements, it is still far from a market leader. Anthropic's Claude Opus 4.8 still leads on every benchmark DeepSeek itself published, while fellow Chinese developer Moonshot still leads among open-weight models with its Kimi K3 model.
Included in Artificial Analysis’ report was a warning that V4-Flash also has a high hallucination rate of 84% among non-correct responses. This means that in cases where it cannot provide the correct response, V4-Flash generally prefers to answer incorrectly rather than declining to answer. The benchmarking firm noted that the newest iteration invented best-guess answers at a lower rate than prior versions, but the rate remains high nonetheless.
Another potential shortcoming is the model’s verbose nature. During testing, it reportedly generated around 206 million output tokens across the entire Intelligence Index test series. This is more than double the median, meaning that the real usage costs of the model can climb higher than its headline fee figures suggest.
