DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp was an experimental multimodal model released on August 21, 2026. DeepSeek retired it on September 10; the legacy ID `deepseek-v4-flash-vision-exp` remains accepted temporarily but now routes to DeepSeek V4.1 Flash. Use `deepseek-flash` for the current hosted multimodal model.
The page remains useful for interpreting August experiments and integrations, but the old ID no longer pins the experimental model. Reproducible current work should use V4.1 Flash and its published model card; historical Vision-Exp results should keep their original date and endpoint context.
DeepSeek's August 21 release documents the original experimental contract. The September 10 release, changelog, and pricing page retire Vision Exp and temporarily alias its ID to V4.1 Flash. Earlier vendor benchmarks and community reports remain historical evidence, not claims about the successor.
- Read screenshots, forms, charts, and diagrams inside an agent workflow.
- Extract or reason over text and structure in images.
- Combine visual observations with function calls or external tools.
- Reuse uploaded images through the Files API.
- Evaluate low-cost multimodal routing before production rollout.
Vision Exp was a separate hosted model, not a vision upgrade to the 0731 checkpoint. DeepSeek has now retired both older Flash models and temporarily maps their legacy IDs to V4.1 Flash. A request that succeeds today under `deepseek-v4-flash-vision-exp` is therefore not evidence that Vision Exp remains live.
DeepSeek accepts JPEG, PNG, GIF, and WebP images as base64 data, public URLs, or reusable Files API IDs. Chat Completions and Anthropic-format Messages accept images only in user messages; Responses carries them as input-image parts. The Files API is useful when the same image is reused or when inline request-size limits would be exceeded.
- Base64 and external-URL images can be up to 32 MiB each; a Files API image can be up to 64 MiB.
- A request can contain up to 600 images within the documented aggregate size limits.
- The `detail: low` option downsizes an image to 512×512 for cheaper processing when fine detail is unnecessary.
DeepSeek resizes images and caps each image at roughly 384 billed tokens. The model uses V4 Flash rates: off peak, $0.007 cache-hit input, $0.22 cache-miss input, and $0.66 output per 1M tokens; peak rates are $0.014, $0.44, and $1.32. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday.
DeepSeek has not published Vision-Exp weights, license metadata for a checkpoint, parameter counts, vision encoder or projector details, training data, a technical report, or a self-hosting guide. Its claim that text capability matches V4 Flash and multimodal-agent performance approaches Opus 4.8 is vendor-reported and should be tested with the same harness and visual tasks before adoption.
Source confidence
DeepSeek API Docs
DeepSeek API Docs
DeepSeek API Docs
DeepSeek API Docs
Reddit r/DeepSeek
DeepSeek V4 Flash Vision Exp FAQ
Common questions about DeepSeek V4 Flash Vision Exp.
Is DeepSeek V4 Flash Vision Exp available now?+
No as a distinct current model. DeepSeek retired Vision Exp; its legacy ID is temporarily accepted but served by V4.1 Flash. Use `deepseek-flash` when you intend the current hosted model.
Can ordinary DeepSeek V4 Flash accept images?+
The retired 0731 checkpoint is text-only. The current V4.1 Flash successor accepts images under `deepseek-flash`; both old hosted IDs temporarily route to that successor.
Can I download DeepSeek V4 Flash Vision Exp weights?+
DeepSeek has not published a Vision-Exp checkpoint or self-hosting guide. The available MIT-licensed `DeepSeek-V4-Flash-0731` weights are for the separate text model and should not be described as Vision-Exp weights.
How much does an image cost?+
Images are converted to input tokens and capped at roughly 384 tokens per image after resizing. Those tokens use the current V4 Flash peak or off-peak input rate, so exact cost depends on image dimensions, cache behavior, time window, and surrounding text.