Signal

DeepSeek releases V4.1 Flash with native vision and lower-cache claims

DeepSeek announced V4.1 Flash on September 10 as the smallest model in its new architecture family, with native multimodal vision, a new asymmetric design, and lower-cost API positioning. The company’s benchmark and efficiency figures remain vendor-reported.

2 min read
DeepSeek official benchmark chart comparing V4.1 Flash with other frontier models
DeepSeek’s official Agentic Benchmark comparison for V4.1 Flash · Credit: DeepSeek View source

DeepSeek published DeepSeek-V4.1 Flash on September 10, presenting it as the smallest model in a new architecture family. The company’s announcement says the model combines native multimodal vision with higher capability, faster inference, greater throughput, and a path to scaling into larger models. Reuters separately reported the launch and those design goals, providing independent confirmation that the release occurred.

An asymmetric architecture

DeepSeek describes V4.1 Flash as a 552-billion-parameter mixture-of-experts model built around a new Causal-Encoder-Decoder structure. Its announcement says only 8 billion input parameters and 16 billion output parameters are activated, a design it presents as a way to lower serving cost relative to models of similar total size. Those architecture and parameter figures come from DeepSeek and have not been independently audited here.

Cache and API changes

The company says the new model reduces KV-cache requirements to one-quarter of the previous model’s HBM demand and one-eighth of its SSD demand. DeepSeek also says V4.1 Flash is available through its API under the name deepseek-flash, with lower pricing taking effect on September 10; older V4 Flash routes are temporarily redirected, and the company plans to route deepseek-v4-pro requests to Flash after September 14 until a future V4.1 Pro release.

DeepSeek labels the announcement’s final section model open source and links a V4.1 Flash model page and technical report on Hugging Face. It says it will support inference adaptations for the open-source community and asks operators with large deployments to bring substantial GPU and storage resources. The release does not independently establish how broadly weights can be deployed, the realized cost reduction, or the reported benchmark advantage.

Sources