Research Obsessions

DeepSeek V4 Flash Vision Exp: $0.08–$0.17 per 1,000 Max-Token Images

DeepSeek's experimental V4 vision endpoint costs $0.084–$0.169 per 1,000 max-token images and trades two benchmark wins with Claude Opus 4.8.

Research brief

What this research adds
This dossier adds a reproducible 1,000-image cost model, a same-file price comparison, and a benchmark-by-benchmark read of DeepSeek's Opus 4.8 claim. It also separates the verified release from five claims the launch material does not support.
Research question
How cheap is DeepSeek-V4-Flash-Vision-Exp for image input, what does its Opus 4.8 comparison prove, and is it ready for production agent workflows?
Method
Reviewed DeepSeek's August 21 announcement, change log, Vision guide, Files API limits, live pricing, model table, rate limits, dsh release notes, and official model repositories. Recalculated image-input costs for 1,000 uploaded 1000×1000 images using each provider's published token rules and standard global API prices. Cross-checked Anthropic and OpenAI pricing against their live documentation. No model API calls were run, so quality findings remain vendor-reported.
Confidence
Medium
Evidence
primary documentation, first-party cost normalization, vendor-reported benchmark analysis
Next verification

TL;DR: DeepSeek caps each image at 384 input tokens. At current V4-Flash rates, 1,000 images at that ceiling cost $0.084 off-peak or $0.169 peak, before text and output tokens. DeepSeek's chart puts it above Opus 4.8 on Agents' Last Exam and ZeroBench, below Opus on ApexBench and Chartography, and 1.4 points below its own text-only V4-Flash baseline on CyberGym. The price is verified; production readiness is not.

Evidence Ledger

ClaimEvidenceSource typeVerifiedConfidenceCaveat
DeepSeek released deepseek-v4-flash-vision-exp through its API platform on August 21, 2026, and labels it experimental.DeepSeek launch announcement and API change logPrimary documentation2026-08-23HighAn API launch is not a production-readiness claim.
The model accepts mixed text and image input through Chat Completions, Messages, and Responses. Images can arrive as base64, external URLs, or file_id references.DeepSeek Vision guidePrimary documentation2026-08-23HighFIM completion is not supported on this model.
DeepSeek resizes images before inference and caps each image at 384 tokens, roughly the token count of an 800×800 image after resizing.DeepSeek Vision token rules and image token calculatorPrimary documentation2026-08-23HighThe API documentation says returned usage is the source of truth because estimates can vary slightly.
V4-Flash image input costs $0.22 per million cache-miss tokens off-peak and $0.44 peak. Therefore 1,000 images at the 384-token ceiling cost $0.08448 off-peak or $0.16896 peak.DeepSeek model pricingPrimary documentation plus first-party calculation2026-08-23HighExcludes prompt text, output, regional terms, taxes, and any cache effect. Peak windows run 01:00–04:00 and 06:00–10:00 UTC on weekdays.
For the same uploaded 1000×1000 file, the published rules produce image-input costs of $2.56 on GPT-5.4, $3.888 on Claude Sonnet 4.6, and $6.48 on Claude Opus 4.8 per 1,000 images.OpenAI image token rules, GPT-5.4 pricing, Claude image tokens, and Claude pricingPrimary documentation plus first-party calculation2026-08-23MediumSame uploaded file does not mean same retained visual detail. DeepSeek downsizes large images more aggressively.
DeepSeek's chart shows V4-Flash-Vision-Exp beating Opus 4.8 on Agents' Last Exam, 27.3 to 25.7, and ZeroBench Pass@5, 35.0 to 34.0.DeepSeek's official benchmark chartVendor-reported benchmark2026-08-23MediumDeepSeek published the run. No independent reproduction was available at verification time.
The same chart shows Opus 4.8 ahead on ApexBench Pass@1, 39.4 to 36.5, and Chartography, 65.0 to 64.3. The multimodal result is a two-two split.DeepSeek's official benchmark chartVendor-reported benchmark2026-08-23MediumFour benchmark scores cannot settle quality across OCR, screenshots, spatial tasks, and long agent loops.
V4-Flash-Vision-Exp scored 75.3 on CyberGym, down from 76.7 for V4-Flash-0731.DeepSeek's benchmark chart and change logVendor-reported benchmark2026-08-23MediumCyberGym measures cybersecurity task capability. It is not a model safety score.
DeepSeek's Files API is free, supports reusable file_id references, permits uploads up to 64 MiB, and allows up to 600 images in one model request.DeepSeek Files API and Vision request limitsPrimary documentation2026-08-23HighReusing a file cuts repeated data transfer. DeepSeek does not say that file_id reuse removes image-token charges.
DeepSeek's dsh app added the model in v0.1.1-rc.1. Version v0.1.1-rc.2 then prioritized Files API uploads and reuse.DeepSeek dsh releasesPrimary release notes2026-08-23HighBoth GitHub tags are prereleases. The announcement shortens the version to “0.1.1.”
DeepSeek has not published a production date or a V4-Flash-Vision-Exp weights repository as of August 23.DeepSeek launch announcement and official Hugging Face organizationPrimary-source absence check2026-08-23MediumThis is a dated negative finding. A repository or date can appear after verification.
V4-Flash-Vision-Exp is not DeepSeek's first image-capable model. DeepSeek-OCR shipped with image understanding and open weights in October 2025.DeepSeek-OCR repository and technical paperPrimary repository and paper2026-08-23HighThe August release is the new image-capable V4 API endpoint, not DeepSeek's first multimodal work.

Where the Evidence Conflicts

The supplied ¥1.15, or roughly $0.17, headline describes the peak ceiling, not the lowest possible price. DeepSeek halves input rates outside two weekday peak windows. The ceiling is $0.084–$0.169 per 1,000.

“Per 1,000 images” is tidy copy and slippery accounting. DeepSeek limits every image to 384 tokens after resizing, while Claude and GPT price images from their own pixel and patch rules. Each model receives that 1000×1000 upload at different detail. DeepSeek wins the bill by a ridiculous margin, while part of that margin buys fewer retained pixels.

DeepSeek is ahead of Opus 4.8 on two multimodal benchmarks and behind on two. Calling that “beats Opus 4.8” is headline tax fraud. My Opus 4.8 builder review applied the same rule to Anthropic: benchmark claims start the test. They do not finish it.

CyberGym does not measure whether a model is safe. It tests cybersecurity capability against real-world vulnerability tasks. The 76.7 to 75.3 change is a 1.4-point capability regression on that evaluation.

I also cut “DeepSeek's first image-capable model.” DeepSeek-OCR predates it by ten months. More precisely, V4-Flash-Vision-Exp adds image input to V4.

I excluded the dsh repository's 30,000-stars-in-one-day claim. GitHub exposes the current total, while dated third-party reports disagree sharply about the first-day count. Popularity is not needed to explain the product change, so the unverifiable number gets no free ride.

What I Tested

I recalculated image-input cost for 1,000 identical 1000×1000 uploads. The calculation uses fresh cache-miss input, standard global endpoints, no text prompt, no output, and prices visible on August 23. I compared identical uploads, not identical internal resolution.

Model and price tierBilled image tokens eachInput price per 1MCost per 1,000 imagesMultiple of DeepSeek peak
DeepSeek V4 Flash Vision Exp, off-peak384 maximum$0.22$0.084480.5×
DeepSeek V4 Flash Vision Exp, peak384 maximum$0.44$0.168961.0×
GPT-5.41,024 patches$2.50$2.5615.2×
Claude Sonnet 51,296$2.00$2.59215.3×
Claude Sonnet 4.61,296$3.00$3.88823.0×
Claude Opus 4.81,296$5.00$6.4838.4×

I used one formula: tokens × 1,000 images × price per token. For GPT-5.4, I counted 32 × 32 patches under OpenAI's published rules for a 1000×1000 image. OpenAI lists no extra multiplier for GPT-5.4, while Anthropic publishes 1,296 tokens for the same uploaded dimensions.

This table corrects another mismatch in the supplied notes. GPT-5.4 costs about $1.92 per 1,000 1024×768 images, but $2.56 for the 1000×1000 files used across this comparison. Claude Sonnet 4.6 lands near $4 because Anthropic bills that square file at 1,296 tokens.

My V4 pricing record shows the previous August rate. DeepSeek changed the table again on August 16. AI prices spoil fast.

Change Log

DateChange foundEvidence affectedConclusion changed?
2026-08-23Initial verification and normalized 1,000-image cost calculationAll claimsInitial draft

My Judgment

My Conclusion

DeepSeek's image-input price is the clear win. Quality is split.

At peak rates, 1,000 max-token images cost $0.169. DeepSeek's own multimodal chart ends two-two against Opus 4.8. A low bill and four vendor-run benchmarks put this model in the test queue, not in charge of a production workflow.

What Builders Should Do

Test it first on high-volume visual work where one missed detail will not ruin the product: screenshot triage, chart extraction, alt-text drafts, interface checks, and visual routing. Keep a stronger fallback for dense documents, tiny text, coordinate-sensitive tasks, and high-stakes extraction.

Log the image-token usage returned by the API. That 384-token cap sets both price and resolution. Your evaluation set should contain the ugliest screenshots your product receives, not the clean launch examples.

What I Would Not Trust Yet

I would not treat the benchmark chart as a production evaluation, or file_id reuse as a token discount. DeepSeek promises only less data transfer for reused files. I would also keep V4-Flash-Vision-Exp away from a workflow that needs a stable model version until DeepSeek publishes a production route, a model card for this variant, or a replacement policy.

What Would Change My Mind

Independent reproductions across OCR, screenshot understanding, chart reading, spatial tasks, and long tool loops would raise my confidence. A production release, weights, guarantees, and sixty days without repricing would help. A first-party test against the exact images a product processes would settle the decision faster than another leaderboard.

FAQ

How Much Does DeepSeek V4 Flash Vision Exp Cost per 1,000 Images?

At the cap, 1,000 images cost $0.08448 off-peak, $0.16896 peak. Those numbers cover image input only, so add prompt text, output tokens, taxes, and any account-specific terms before forecasting a real workload. Smaller images may use fewer tokens, although DeepSeek scales very small images up before inference. Log the API's returned usage during your test rather than treating 384 as every image's guaranteed token count.

Is DeepSeek V4 Flash Vision Exp Cheaper Than Claude and GPT for Image Input?

Yes, by a large margin under the published rules. For 1,000 uploaded 1000×1000 images, this dossier calculates $0.084–$0.169 on DeepSeek, $2.56 on GPT-5.4, $2.592 on Claude Sonnet 5, $3.888 on Sonnet 4.6, and $6.48 on Opus 4.8. The comparison has one important limit: DeepSeek downsizes large inputs to roughly an 800×800 pixel equivalent. Their retained visual detail still differs.

Does DeepSeek V4 Flash Vision Exp Beat Claude Opus 4.8?

DeepSeek's chart ends in a two-two benchmark draw. V4-Flash-Vision-Exp leads Agents' Last Exam by 1.6 points and ZeroBench Pass@5 by 1.0, while Opus 4.8 leads ApexBench Pass@1 by 2.9 points and Chartography by 0.7. That supports “competitive with Opus 4.8” for the tested tasks, not a general quality win across OCR, UI screenshots, documents, or production agent loops.

DeepSeek ran and published the evaluation. No independent reproduction was available on August 23.

Is DeepSeek V4 Flash Vision Exp Open Source or Production Ready?

The API works now, but DeepSeek calls it experimental. The August 21 announcement provides no production date, dedicated weights repository, or variant-specific model card, although DeepSeek has open-sourced other V4 weights and earlier vision work. An eventual release remains plausible and unpublished.

Use it for reversible evaluations. Wait for version guarantees and a private regression suite before assigning legally, financially, or physically risky work.

What Does DeepSeek's Free Files API Change?

The Files API lets you upload an image once and reference it later by file_id, cutting repeated request payloads for workflows that reuse the same screenshot, document page, or chart. Uploads can be as large as 64 MiB, a model request can contain up to 600 images, and each user gets 25 GiB across at most 10,000 stored files.

“Free” refers to file storage and transfer. The image still enters model context. DeepSeek never promises that reused files escape image-token billing.

Share this with a friend who is pricing a screenshot-heavy agent before the image bill arrives.

Subscribe to Product with Attitude for more free AI model cost breakdowns for builders choosing what belongs in production.