DeepSeek V4.1 Flash
Fast million-token reasoning model for coding, agents, and high-throughput workflows.
Fast million-token reasoning model for coding, agents, and high-throughput workflows. This is the recommended default model.
| Model ID | deepseek-v4.1-flash |
| Context window | 1M tokens |
| Max output | 384K tokens |
| Modalities | text, image |
| Alias | zro/deepseek-v4.1-flash |
Pricing
DeepSeek V4.1 Flash pricing: $0.15 per 1M input tokens, $0.60 per 1M output tokens, and $0.003 per 1M cache read tokens while the 50% promotional rate is active. List prices are $0.30 input, $1.20 output, and $0.006 cache read.
| Input | Output | Cache read |
|---|---|---|
$0.15 (list $0.30) | $0.60 (list $1.20) | $0.003 (list $0.006) |
Reasoning effort
Default: high. Available levels: low, high, max.
Upstream weights
deepseek-ai/DeepSeek-V4.1-Flash (MIT License).