DeepSeek V4 Pro
Current DeepSeek pricing and changelog list continued V4 Pro service. Ready to experience the power of AI? Start your journey here!
Platform: Replicate
Language ModelReasoningTool CallingLong Context
DeepSeek API
Off-peak: $0.022/1M cache-hit input, $0.66/1M cache-miss input, $1.98/1M output; peak: $0.044, $1.32, and $3.96 respectively
Commercial🚀Function Overview
Current pricing and changelog agree that V4 Pro service continues with unchanged billing after September 14. See what makes this AI model special!
Key Features
- 1M token context length in the official DeepSeek model table
- 384K maximum output in the official DeepSeek model table
- Low, high, and max reasoning effort plus non-thinking mode switching guidance
- JSON output and tool calls
- Chat prefix completion beta and FIM completion in non-thinking mode
- OpenAI-format and Anthropic-format API base URLs
- Responses API support in the current official model table
- Current pricing and changelog retain V4-Pro-0813; older September 10 redirect text remains the conflicting source.
- Generally available on DeepSeek app, web, and API since August 13, 2026
- Off-peak pricing of $0.022 cache-hit input, $0.66 cache-miss input, and $1.98 output per 1M tokens
- Peak pricing of $0.044 cache-hit input, $1.32 cache-miss input, and $3.96 output per 1M tokens
- No separately identified 0813 public checkpoint as of August 23, 2026
- October 7 check: current pricing and changelog list continued V4-Pro-0813 service; the older September 10 launch article still contains a redirect claim
Use Cases
- •Long-context chat and agent workflows
- •Coding-agent sessions that need tool calls and cache-aware cost control
- •Structured JSON output generation
- •API experiments comparing DeepSeek pricing against other frontier or long-context models
⚙️Input Parameters
messages
arrayChat messages sent to the DeepSeek API using OpenAI-format or Anthropic-format compatible endpoints.
thinking_mode
stringDeepSeek documents support for both thinking and non-thinking modes, with thinking enabled by default.
tools
arrayOptional tool definitions for tool-calling workflows.
response_format
objectOptional JSON output controls when structured output is needed.
💡Usage Examples
Example 1
Input Parameters
{
"model": "deepseek-v4-pro",
"messages": [
{
"role": "user",
"content": "Summarize the tradeoffs of cached input pricing for a long coding-agent session."
}
],
"thinking_mode": "default"
}Output Results
A chat completion response from the DeepSeek API. Verify the current request and response schema in the official DeepSeek API documentation before production use.
Quick Actions
Technical Specifications
- Hardware Type
- DeepSeek API
- Commercial Use
- Supported
- Pricing
- Off-peak: $0.022/1M cache-hit input, $0.66/1M cache-miss input, $1.98/1M output; peak: $0.044, $1.32, and $3.96 respectively
- Platform
- Replicate