DeepSeek V4.1 Flash Launches With Lower Prices and Native Vision
DeepSeek has officially launched V4.1 Flash, replacing its previous V4 Flash and V4 Flash Vision Experimental models while cutting API prices and adding native image understanding.
The new model became available on September 10, 2026, under the API name deepseek-flash. Older Flash model names will continue to work as aliases but now route requests to V4.1 Flash.
Faster Architecture and Native Vision
V4.1 Flash uses a redesigned architecture with around 748 billion total parameters, including a 552B backbone and 196B Engram parameters.
Despite its size, only around 8B parameters per token are active during prefill and 16B during decoding, helping keep inference costs lower.
The model also introduces native multimodal capabilities. Its DeepSeek-ViT encoder can process images alongside text, with support for resolutions of up to roughly 1344 × 1344 pixels.
V4.1 Flash supports a context window of up to 1 million tokens and output lengths of up to 384,000 tokens.
Developers can also adjust reasoning effort from 1 to 100, allowing them to balance cost and accuracy.
| Specification | DeepSeek V4.1 Flash |
|---|---|
| Release Date | September 10, 2026 |
| API Model Name | deepseek-flash |
| Total Parameters | Approx. 748B |
| Active Parameters | ~8B during prefill, ~16B during decode |
| Architecture | Mixture-of-Experts, Causal Encoder-Decoder |
| Context Window | Up to 1 million tokens |
| Maximum Output | Up to 384K tokens |
| Vision Support | Native multimodal image understanding |
| Attention System | Compressed Sparse Attention 2 |
| Reasoning Control | Adjustable from 1–100 |
| Training Data | 45 trillion multimodal tokens |
| License | MIT, open weights |
| Models Replaced | V4 Flash, V4 Flash Vision Exp |
| Next Model | V4.1 Pro planned |
Stronger Agent and Coding Performance
DeepSeek’s own benchmarks show major improvements over V4 Flash and, in several areas, V4 Pro.
Its DeepSWE score increased from 54.4 on V4 Flash to 74.2, while Terminal Bench 2.1 rose from 82.7 to 90.6.
V4.1 Flash also scored 88.1 on CyberGym and 54.8 on AutomationBench.
Independent testing has also been positive. Artificial Analysis reported a 68.9% score on AutomationBench-AA, while Vals.ai gave the model a 57.86% Vals Index, placing it at the top among open-weight models in that comparison.
However, V4.1 Flash does not lead every benchmark. It trails some frontier models on newer Terminal Bench tests and performs poorly on certain specialized evaluations such as SRE Bench and Harvey’s Legal Agent benchmark.
Lower API Prices
DeepSeek has also reduced Flash pricing significantly.
During off-peak hours, cached input now costs $0.003 per million tokens, uncached input costs $0.15, and output costs $0.60.
Peak pricing doubles those figures to $0.006 for cached input, $0.30 for uncached input, and $1.20 for output per million tokens.
The biggest reduction applies to cached input, which should particularly benefit agent workloads that repeatedly reuse large context windows.
V4 Pro Is Also Being Retired
DeepSeek is also preparing to phase out V4 Pro.
Starting September 14 at 04:00 UTC, requests sent to deepseek-v4-pro will automatically be redirected to V4.1 Flash and charged at the new Flash rates.
That arrangement will remain in place until V4.1 Pro becomes available.
The migration could substantially reduce costs for existing V4 Pro users, although the new architecture may behave differently in areas such as prompting, tool calling, and response style.
The post DeepSeek V4.1 Flash Launches With Lower Prices and Native Vision appeared first on ProPakistani.



