All changelog entries
FeatureNew0.14.0ModelsVision
DeepSeek-V4.1-Flash — a big step up for agents and terminals
September 18, 2026
DeepSeek-V4.1-Flash is live. It reads text and images, carries a 1M-token context window, and is built for input-heavy agentic and coding work. Reasoning effort is adjustable, so it can answer quickly or think hard on request. Call it as DeepSeek-V4.1-Flash.
DeepSeek publishes these scores against the earlier V4 Flash, all at maximum reasoning effort:
| Benchmark | V4.1 Flash | V4 Flash |
|---|---|---|
| Terminal-Bench 2.1 | 90.6 | 82.7 |
| Terminal-Bench 3.0 | 30.0 | 7.6 |
| DeepSWE v1.1 (resolved) | 74.2 | 54.4 |
| NL2Repo-Bench | 64.0 | 54.2 |
| SEC-Bench Pro | 62.8 | 30.9 |
| HLE with tools | 63.9 | 51.5 |
| Codeforces rating | 3471 | 3289 |
| GPQA Diamond | 90.9 | 89.9 |
It costs 2 requests of quota per call and is available on every paid plan and the free tier. DeepSeek-V4-Flash stays exactly as it is at 1 request per call, so nothing changes for existing setups. Pick V4.1 when the task is a long agent run or terminal work, and V4 when you want the cheapest fast model.
Source: DeepSeek-V4.1-Flash model card.
