AI KNOWLEDGE DESK

Models · entities · concepts · comparisons · practical tools

GETLLMS.ORG
ModelMultimodal AI models

DeepSeek V4 Flash Vision Exp

DeepSeek V4 Flash Vision Exp was an experimental multimodal model released on August 21, 2026. DeepSeek retired it on September 10; the legacy ID `deepseek-v4-flash-vision-exp` remains accepted temporarily but now routes to DeepSeek V4.1 Flash. Use `deepseek-flash` for the current hosted multimodal model.

Why it matters

The page remains useful for interpreting August experiments and integrations, but the old ID no longer pins the experimental model. Reproducible current work should use V4.1 Flash and its published model card; historical Vision-Exp results should keep their original date and endpoint context.

Source-backed summary

DeepSeek's August 21 release documents the original experimental contract. The September 10 release, changelog, and pricing page retire Vision Exp and temporarily alias its ID to V4.1 Flash. Earlier vendor benchmarks and community reports remain historical evidence, not claims about the successor.

Primary use cases
  • Read screenshots, forms, charts, and diagrams inside an agent workflow.
  • Extract or reason over text and structure in images.
  • Combine visual observations with function calls or external tools.
  • Reuse uploaded images through the Files API.
  • Evaluate low-cost multimodal routing before production rollout.
The legacy ID now names a compatibility route

Vision Exp was a separate hosted model, not a vision upgrade to the 0731 checkpoint. DeepSeek has now retired both older Flash models and temporarily maps their legacy IDs to V4.1 Flash. A request that succeeds today under `deepseek-v4-flash-vision-exp` is therefore not evidence that Vision Exp remains live.

How image input works

DeepSeek accepts JPEG, PNG, GIF, and WebP images as base64 data, public URLs, or reusable Files API IDs. Chat Completions and Anthropic-format Messages accept images only in user messages; Responses carries them as input-image parts. The Files API is useful when the same image is reused or when inline request-size limits would be exceeded.

  • Base64 and external-URL images can be up to 32 MiB each; a Files API image can be up to 64 MiB.
  • A request can contain up to 600 images within the documented aggregate size limits.
  • The `detail: low` option downsizes an image to 512×512 for cheaper processing when fine detail is unnecessary.
Image tokens and current pricing

DeepSeek resizes images and caps each image at roughly 384 billed tokens. The model uses V4 Flash rates: off peak, $0.007 cache-hit input, $0.22 cache-miss input, and $0.66 output per 1M tokens; peak rates are $0.014, $0.44, and $1.32. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday.

What remains unknown

DeepSeek has not published Vision-Exp weights, license metadata for a checkpoint, parameter counts, vision encoder or projector details, training data, a technical report, or a self-hosting guide. Its claim that text capability matches V4 Flash and multimodal-agent performance approaches Opus 4.8 is vendor-reported and should be tested with the same harness and visual tasks before adoption.

DeepSeek V4 Flash Vision Exp FAQ

Common questions about DeepSeek V4 Flash Vision Exp.

Is DeepSeek V4 Flash Vision Exp available now?+

No as a distinct current model. DeepSeek retired Vision Exp; its legacy ID is temporarily accepted but served by V4.1 Flash. Use `deepseek-flash` when you intend the current hosted model.

Can ordinary DeepSeek V4 Flash accept images?+

The retired 0731 checkpoint is text-only. The current V4.1 Flash successor accepts images under `deepseek-flash`; both old hosted IDs temporarily route to that successor.

Can I download DeepSeek V4 Flash Vision Exp weights?+

DeepSeek has not published a Vision-Exp checkpoint or self-hosting guide. The available MIT-licensed `DeepSeek-V4-Flash-0731` weights are for the separate text model and should not be described as Vision-Exp weights.

How much does an image cost?+

Images are converted to input tokens and capped at roughly 384 tokens per image after resizing. Those tokens use the current V4 Flash peak or off-peak input rate, so exact cost depends on image dimensions, cache behavior, time window, and surrounding text.