DeepSeek on Friday released an experimental vision model that the Chinese AI startup said brings its multimodal agent performance close to Anthropic's Opus-4.8 on agent benchmarks.

The model, DeepSeek-V4-Flash-Vision-Exp, is now available through the DeepSeek API Platform, the company said in an announcement. The model matches V4 Flash on text capabilities, including agents, reasoning, and world knowledge, while adding image inputs, it added.

The company said the new model made a “major leap” over V4 Flash on multimodal agent benchmarks, bringing its performance close to Opus-4.8. DeepSeek did not provide benchmark scores in its announcement.

Notably, the model can analyze images alongside text, including describing pictures, reading text from screenshots, and analyzing charts. It supports JPEG, PNG, GIF, and WebP images.

DeepSeek said V4 Flash Vision supports Chat Completions, Messages, and Responses APIs. Images can be accepted through base64 data, external URLs, or its newly launched Files API.

Images are tokenized for billing, with each image consuming up to 384 tokens at V4 Flash pricing. The model can process up to 600 images in a single request, while individual images can be as large as 32 megabytes through base64 or external URLs and 64 megabytes when referenced through the Files API.

The Files API allows users to upload an image once and reference it through a file ID across requests. DeepSeek said the service is free to use and can also reduce bandwidth when the same image is reused.

The release comes as DeepSeek has resumed its second funding round, seeking close to $8 billion, with Monolith Management in talks to participate, Bloomberg reported earlier this month. The company is seeking a valuation close to 500 billion yuan ($74.3 billion).