The GenAI Field Guide · Trending

DeepSeek V4.1-Flash

An MIT-licensed, 552B-parameter mixture-of-experts model that reads text and images, takes a one million token context, and runs on DeepSeek's API for cents per million tokens.

DeepSeek V4.1-Flash is the first model in a new DeepSeek architecture that splits reading and writing into different halves of the network and shrinks the KV cache to about a quarter of the previous generation. DeepSeek says it beats its own V4-Pro on agentic and coding work, and it now serves V4-Pro traffic at Flash prices. Anyone choosing a cheap default model, or anyone with a production integration pointed at a DeepSeek endpoint, should look at it now.

What it is

DeepSeek released V4.1-Flash on 10 September 2026, publishing the weights on Hugging Face under the MIT license alongside a technical paper. On the API it is the model called deepseek-flash. DeepSeek calls it the smallest model in its new architecture family; Reuters and The Standard carried that framing, and Reuters tied the timing to DeepSeek's preparation for a listing on Shanghai's STAR Market.

Under the name, it is not small. The backbone has 552 billion parameters, it accepts images natively, and the context window is one million tokens. What makes it a Flash model is how little of it runs per token: about 8 billion parameters while reading the prompt and about 16 billion while generating. Some developers have questioned whether a 552B model deserves the Flash label at all, as Genuine Impact reported.

The commercial move matters as much as the model. DeepSeek's release note says requests to V4-Pro are redirected to V4.1-Flash pricing from 14 September 2026 at 04:00 UTC, and Bloomberg, as carried by Deccan Chronicle, reported that all V4-Pro inference is rerouted to the new model. Bloomberg also reported that DeepSeek says it outperforms Moonshot's Kimi K3 and V4-Pro on coding and agentic tasks while still trailing the flagship models from Anthropic and OpenAI.

How it works

DeepSeek calls the design a Causal Encoder-Decoder. The model card describes a 40-layer Transformer split into a 20-layer encoder and a 20-layer decoder. The encoder handles prefill, the step where the model reads your prompt, and activates about 8B parameters per token; the decoder handles generation and activates about 16B. Each mixture-of-experts layer has one shared expert and 384 routed experts, with six routed experts active per token. Reading is usually the larger share of tokens in agent and retrieval workloads, so putting the cheaper half there is where the savings come from.

The second idea is KV cache compression. The model card gives the global KV cache as about 890 bytes per token, roughly a quarter of DeepSeek V4-Flash, stored in a 4-bit floating point format with a shared scale per 16 channels. It also describes a second version of DeepSeek's compressed sparse attention, with hierarchical sparse indexing. DeepSeek's release note says the cache needs a quarter of the GPU memory and an eighth of the SSD storage of the previous generation. A smaller cache is what makes a one million token context and cheap cached input prices practical to serve.

Vision comes from a DeepSeek ViT encoder built into the model, so the separate vision preview models are retired. Training used a 45 trillion token multimodal corpus, with sparse attention trained at 64K tokens and extended to one million. Reasoning effort is continuously controllable from 1 to 100 according to the model card, and the published evaluations use the maximum. The API exposes this through a thinking switch and a reasoning_effort parameter.

How to use it

Anyone with a DeepSeek platform account and API key can call it today through an OpenAI-compatible endpoint, and the API also lists Anthropic-format and Responses-style compatibility. The weights are on Hugging Face if you would rather run it yourself, but at 552B parameters that means a multi-GPU server; the model card does not state a GPU count.

  1. Create an API key on the DeepSeek platform and set it as DEEPSEEK_API_KEY.
  2. Point any OpenAI SDK at the base URL api.deepseek.com and use the model name deepseek-flash. The old names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work but are served by V4.1-Flash.
  3. Turn thinking on or off per request with the thinking field and set reasoning_effort; start low for extraction and classification and raise it only where your evals show a gain.
  4. If you run batch or background jobs, schedule them outside the peak windows (01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday), when prices are half.
  5. Keep a stable prefix at the top of long prompts so repeat calls hit the cache, which is priced at a small fraction of fresh input.
  6. To self-host, serve deepseek-ai/DeepSeek-V4.1-Flash with vLLM or SGLang as the model card shows, and use DeepSeek's Python encoding reference or its deepseek-recipe library for prompt formatting, because the release ships no Jinja chat template.

Use cases

Sources

  1. DeepSeek-V4.1-Flash release note, DeepSeek API Docs, 2026-09-10
  2. Models and pricing, DeepSeek API Docs, 2026-10-02
  3. Your first API call, DeepSeek API Docs, 2026-10-02
  4. deepseek-ai/DeepSeek-V4.1-Flash model card, Hugging Face, 2026-09-10
  5. China's DeepSeek launches V4.1-Flash model (Reuters), The Business Standard, 2026-09-10
  6. China's DeepSeek launches V4.1-Flash model, The Standard, 2026-09-10
  7. China's Deepseek Unveils V4.1-Flash Model (Bloomberg), Deccan Chronicle, 2026-09-10
  8. Inside DeepSeek V4.1 Flash: A Cheaper Model, a Bigger Signal, Genuine Impact, 2026-09-15
  9. DeepSeek-V4.1-Flash Benchmarks, Pricing and Context Window, llm-stats, 2026-10-02