DeepSeek launches ultra-cheap vision model at 1 yuan per 1,000 images, outpaces Claude Opus on key benchmarks

Visual understanding shifts from premium add-on to baseline capability
DeepSeek's aggressive pricing strategy on its new vision model transforms how developers think about using images in AI agent workflows.
Mark

Why release a vision model now, and why attach it to V4-Flash instead of the more powerful V4-Pro?

Mimi

The community had already built workarounds—they were layering other vision models on top of DeepSeek's text engine because they needed it. But more importantly, the target use case isn't casual image chat. It's agents reading screenshots, analyzing interfaces, inspecting designs. Those tasks need speed and cost efficiency, not maximum precision on a single image. V4-Flash is fast and cheap. That's the point.

Mark

The pricing is striking—seventeen cents per thousand images. How does that actually work technically?

Mimi

Every image gets resized to an 800-by-800 equivalent before processing, and the model caps token consumption at 384 tokens per image, no matter how detailed the original. That's the lever. Competitors charge based on actual token consumption, which climbs with resolution. DeepSeek just stops it from climbing.

Mark

Does that mean the model is less capable at understanding high-resolution images?

Mimi

It's a trade-off. The model sees less pixel-level detail, but for the workflows DeepSeek is targeting—reading text on a webpage, spotting an error in a UI, analyzing a chart—that's acceptable. And the cost difference is enormous. Developers can now afford to take a screenshot at every step of an agent workflow instead of rationing vision calls.

Mark

The model surpassed Claude Opus on some benchmarks. Is this a production-ready system or still experimental?

Mimi

It's labeled experimental, and DeepSeek hasn't released the weights or a technical report yet. But the company has a pattern: V3.2-Exp went through public testing before becoming production. Given that V4-Flash moved from preview to general availability in three months, the vision version could follow the same path. For now, enterprise users have to accept some uncertainty.

Mark

What does this mean for the broader AI market?

Mimi

It's pressure. Competitors charge based on actual token consumption per image. DeepSeek just made visual understanding cheap enough that it becomes a default operation in agent workflows, not a carefully budgeted feature. That changes the economics of what's possible.

  • DeepSeek's V4-Flash-Vision-Exp processes 1,000 images for roughly seventeen cents — ten to twenty times cheaper than Claude, GPT, and Gemini equivalents — creating immediate pricing pressure across the industry.
  • The release fills a critical gap exposed two weeks earlier when DeepSeek's own Harness agent framework went viral but couldn't handle images, forcing developers into awkward workarounds with third-party vision tools.
  • Despite its budget positioning, the model outperforms Anthropic's flagship Claude Opus 4.8 on two key agent benchmarks, challenging the assumption that low cost means lower capability.
  • A simultaneously released Files API — currently free — lets developers upload images once and reuse them across hundreds of requests, further compressing the operational cost of vision-heavy agent workflows.
  • The 'Exp' experimental label and absent technical report suggest this is a validation step, with a production release likely following the same three-month pattern DeepSeek used for its V4-Flash model.

On August 21, DeepSeek released its first vision-capable model, pricing image processing at a fraction of what industry leaders charge — a move that quietly redraws the economics of AI agent development. By capping token consumption per image and targeting high-throughput workflows rather than casual use, the company is not merely competing on price but proposing a different philosophy: that visual understanding should be a default capability, not a rationed luxury. Released alongside a native update to its open-source agent framework, the launch signals DeepSeek's ambition to build not just a model, but an entire infrastructure for the next generation of autonomous software.

On August 21, DeepSeek released its first image-processing model — DeepSeek-V4-Flash-Vision-Exp — opening it immediately to developers via API. The pricing was the headline: roughly seventeen cents to process a thousand images, dropping below six cents during off-peak hours. By comparison, Anthropic's Claude Sonnet charges around four dollars for the same workload, OpenAI's GPT-5.4 Vision nearly two dollars, and Google's Gemini about fifty cents. DeepSeek achieves its price by automatically resizing every image to an 800-by-800 pixel equivalent and capping token consumption at 384 tokens per image — a hard ceiling that keeps costs flat regardless of source resolution.

The timing was deliberate. Two weeks earlier, DeepSeek had open-sourced Harness, an agent-building framework that accumulated thirty thousand GitHub stars in a single day. But Harness exposed an embarrassing gap: DeepSeek's existing models couldn't process images at all. Developers built workarounds — layering third-party vision tools or using OCR as a substitute. V4-Flash-Vision-Exp closed that gap directly, and on the same day, Harness released version 0.1.1 with native support for the new model.

On performance, the model held its ground. It matched V4-Flash on text reasoning while maintaining a one-million-token context window. More strikingly, it surpassed Claude Opus 4.8 on two multimodal agent benchmarks — scoring 27.3 versus 25.7 on an agent reasoning exam, and 35.0 versus 34.0 on an image-reading test. A safety benchmark showed a minor decline. DeepSeek illustrated the model's intended use through case studies involving multi-step agent workflows: generating business presentations from visual inputs, building interactive web elements through iterative conversation, and analyzing interface designs — tasks where speed and cost matter more than perfect single-image comprehension.

A companion Files API, currently free, allows developers to upload images once and reference them across multiple requests, with a single call capable of handling up to six hundred images. The model's weights and technical report have not yet been published, and the 'Exp' designation suggests an experimental phase before a formal production release — a pattern DeepSeek followed with V4-Flash, which reached general availability roughly three months after its preview.

The strategic logic is legible: by attaching vision to V4-Flash rather than its more powerful V4-Pro, DeepSeek is explicitly targeting cost-sensitive, high-volume agent work. The economics now make it practical to take a screenshot at every step of an agent workflow rather than carefully rationing vision calls. Combined with Harness and the Files API, DeepSeek is assembling a full-stack infrastructure for agent development — and leaving competitors to decide whether to match the pricing or absorb the pressure.

On August 21, DeepSeek released its first image-processing model, DeepSeek-V4-Flash-Vision-Exp, opening it immediately to developers through its API platform. The move was not subtle. The company priced it aggressively: roughly 1.15 yuan—about seventeen cents—to process a thousand images at standard resolution. During off-peak hours, that drops below six cents per thousand.

To understand what this means, consider what the rest of the industry charges for the same work. Anthropic's Claude Sonnet 4.6 costs around four dollars per thousand images at the same resolution. OpenAI's GPT-5.4 Vision runs nearly two dollars. Google's Gemini 3.1 Pro comes in at about fifty cents. DeepSeek's approach achieves its price through a technical choice: every image, regardless of how detailed the original, gets automatically resized to an 800-by-800 pixel equivalent before processing. The model then caps token consumption at 384 tokens per image, a hard ceiling that prevents costs from climbing with resolution. The result is a pricing structure that undercuts competitors by a factor of ten to twenty.

The timing mattered. Two weeks earlier, DeepSeek had open-sourced Harness, a framework for building AI agents, under the MIT license. It accumulated thirty thousand GitHub stars in its first day. But Harness revealed a problem: DeepSeek's existing V4 models couldn't process images at all. Developers who tried to upload images received an error message. The community responded by building workarounds—some layered third-party vision models on top of DeepSeek's text engine, others used optical character recognition and pixel analysis as indirect paths to image understanding. V4-Flash-Vision-Exp filled that gap directly, and on the same day, Harness released version 0.1.1 with native support for the new model.

On performance, the model held its own. On pure text tasks—reasoning, agent behavior, world knowledge—it matched V4-Flash's existing capabilities while maintaining a context window of one million tokens and maximum output of 384,000 tokens. More notably, on four multimodal agent benchmarks, it surpassed Anthropic's flagship Claude Opus 4.8 on two measures: scoring 27.3 versus 25.7 on an agent-oriented exam, and 35.0 versus 34.0 on an image-reading test called ZeroBench. A safety benchmark called Cybergym showed a slight decline, dropping from 76.7 to 75.3.

DeepSeek demonstrated the model's intended use through three case studies: generating a business presentation for a Tibet self-driving tour targeting wealthy clients, building interactive website elements through multiple turns of conversation, and creating a front-end demo with dynamic visual effects. The common thread was not casual image questions but workflows where visual input fed into multi-step agent operations—reading screenshots, identifying interface errors, analyzing charts, inspecting software designs. These tasks demand speed and cost efficiency more than they demand perfect single-image understanding.

The company also released a Files API, currently free, that lets developers upload images once and reference them across multiple requests without re-uploading. A single request can process up to six hundred images. Three input methods are supported: Base64-encoded local images, external URLs, and the new file-upload system.

DeepSeek has not yet released the model's weights, published a technical report, or shared detailed vision benchmarks. The "Exp" designation—experimental—suggests the company may be following its earlier pattern with V3.2-Exp: validating a new approach with a public experimental version, gathering feedback, then releasing a production version. The V4-Flash production model moved from preview to general availability in roughly three months, so the vision version's formalization may follow a similar timeline.

The strategic signal is clear. By bringing vision to V4-Flash rather than the more powerful V4-Pro, DeepSeek is targeting high-throughput, cost-sensitive agent work. The pricing structure transforms visual understanding from a premium add-on into a baseline capability. For developers building agents, the economics now favor taking a screenshot at every step rather than carefully rationing vision calls. Combined with Harness's plugin architecture and the Files API's efficiency, DeepSeek is assembling a complete infrastructure for agent work—from the model layer down through the toolchain. The question now is whether competitors will follow, or whether this pricing pressure forces a reckoning across the industry.

The target use case is not casual image chat, but high-throughput, low-cost agent vision tasks—reading webpage screenshots, identifying error interfaces, analyzing charts, inspecting software UIs.
— DeepSeek product strategy (inferred from case studies and pricing structure)
Contáctanos FAQ