DeepSeek released its V4.1-Flash model on Thursday, introducing a 552B-parameter mixture-of-experts architecture with 8B active parameters for input and 16B for output.
The model uses a new Causal Encoder-Decoder architecture and supports native multimodal inputs, including visual understanding, according to a statement. DeepSeek said V4.1-Flash is the smallest model in its new architecture family.
A key change is the model’s reduced KV cache footprint, the company noted, adding that V4.1-Flash requires one-quarter the HBM and one-eighth the SSD storage used by the previous generation.
According to the company, the smaller cache reduces storage requirements for agentic workloads, where cache-hit costs can account for a significant share of overall usage costs.
V4.1-Flash is now available through DeepSeek’s API, while the older V4-Flash and V4-Flash-Vision-Exp models have been discontinued. Their API names will temporarily route requests to V4.1-Flash for compatibility.
Notably, V4.1-Flash surpassed V4-Pro on performance, cost, speed, and total runtime in tests by multiple parties, and the company is therefore phasing out the latter.
Beginning next week on Sept. 14 at 04:00 UTC, DeepSeek will route all requests made to V4 Pro will be routed to V4.1-Flash and billed at the latest model's rates. The arrangement will remain in place until DeepSeek launches V4.1-Pro, per the statement.
The company has adjusted V4.1-Flash API pricing, with off-peak rates set at half the peak price. The new pricing took effect at 04:00 UTC on Sept. 10. Tencent's WorkBuddy and CodeBuddy, along with OpenCode, now support the model.
