DeepSeek published DeepSeek-V4.1 Flash on September 10, presenting it as the smallest model in a new architecture family. The company’s announcement says the model combines native multimodal vision with higher capability, faster inference, greater throughput, and a path to scaling into larger models. Reuters separately reported the launch and those design goals, providing independent confirmation that the release occurred.
An asymmetric architecture
DeepSeek describes V4.1 Flash as a 552-billion-parameter mixture-of-experts model built around a new Causal-Encoder-Decoder structure. Its announcement says only 8 billion input parameters and 16 billion output parameters are activated, a design it presents as a way to lower serving cost relative to models of similar total size. Those architecture and parameter figures come from DeepSeek and have not been independently audited here.
Cache and API changes
The company says the new model reduces KV-cache requirements to one-quarter of the previous model’s HBM demand and one-eighth of its SSD demand. DeepSeek also says V4.1 Flash is available through its API under the name deepseek-flash, with lower pricing taking effect on September 10; older V4 Flash routes are temporarily redirected, and the company plans to route deepseek-v4-pro requests to Flash after September 14 until a future V4.1 Pro release.
DeepSeek labels the announcement’s final section model open source and links a V4.1 Flash model page and technical report on Hugging Face. It says it will support inference adaptations for the open-source community and asks operators with large deployments to bring substantial GPU and storage resources. The release does not independently establish how broadly weights can be deployed, the realized cost reduction, or the reported benchmark advantage.
