The GenAI Field Guide · Trending
An MIT-licensed, 552B-parameter mixture-of-experts model that reads text and images, takes a one million token context, and runs on DeepSeek's API for cents per million tokens.
DeepSeek V4.1-Flash is the first model in a new DeepSeek architecture that splits reading and writing into different halves of the network and shrinks the KV cache to about a quarter of the previous generation. DeepSeek says it beats its own V4-Pro on agentic and coding work, and it now serves V4-Pro traffic at Flash prices. Anyone choosing a cheap default model, or anyone with a production integration pointed at a DeepSeek endpoint, should look at it now.
DeepSeek released V4.1-Flash on 10 September 2026, publishing the weights on Hugging Face under the MIT license alongside a technical paper. On the API it is the model called deepseek-flash. DeepSeek calls it the smallest model in its new architecture family; Reuters and The Standard carried that framing, and Reuters tied the timing to DeepSeek's preparation for a listing on Shanghai's STAR Market.
Under the name, it is not small. The backbone has 552 billion parameters, it accepts images natively, and the context window is one million tokens. What makes it a Flash model is how little of it runs per token: about 8 billion parameters while reading the prompt and about 16 billion while generating. Some developers have questioned whether a 552B model deserves the Flash label at all, as Genuine Impact reported.
The commercial move matters as much as the model. DeepSeek's release note says requests to V4-Pro are redirected to V4.1-Flash pricing from 14 September 2026 at 04:00 UTC, and Bloomberg, as carried by Deccan Chronicle, reported that all V4-Pro inference is rerouted to the new model. Bloomberg also reported that DeepSeek says it outperforms Moonshot's Kimi K3 and V4-Pro on coding and agentic tasks while still trailing the flagship models from Anthropic and OpenAI.
DeepSeek calls the design a Causal Encoder-Decoder. The model card describes a 40-layer Transformer split into a 20-layer encoder and a 20-layer decoder. The encoder handles prefill, the step where the model reads your prompt, and activates about 8B parameters per token; the decoder handles generation and activates about 16B. Each mixture-of-experts layer has one shared expert and 384 routed experts, with six routed experts active per token. Reading is usually the larger share of tokens in agent and retrieval workloads, so putting the cheaper half there is where the savings come from.
The second idea is KV cache compression. The model card gives the global KV cache as about 890 bytes per token, roughly a quarter of DeepSeek V4-Flash, stored in a 4-bit floating point format with a shared scale per 16 channels. It also describes a second version of DeepSeek's compressed sparse attention, with hierarchical sparse indexing. DeepSeek's release note says the cache needs a quarter of the GPU memory and an eighth of the SSD storage of the previous generation. A smaller cache is what makes a one million token context and cheap cached input prices practical to serve.
Vision comes from a DeepSeek ViT encoder built into the model, so the separate vision preview models are retired. Training used a 45 trillion token multimodal corpus, with sparse attention trained at 64K tokens and extended to one million. Reasoning effort is continuously controllable from 1 to 100 according to the model card, and the published evaluations use the maximum. The API exposes this through a thinking switch and a reasoning_effort parameter.
Anyone with a DeepSeek platform account and API key can call it today through an OpenAI-compatible endpoint, and the API also lists Anthropic-format and Responses-style compatibility. The weights are on Hugging Face if you would rather run it yourself, but at 552B parameters that means a multi-GPU server; the model card does not state a GPU count.