DeepSeek released V4.1-Flash on September 10, 2026, and the version number undersells it. This is not a point release of V4-Flash. It is the first model on a new architecture family the company calls Causal Encoder-Decoder (CED): a 552-billion-parameter Mixture-of-Experts model that deliberately breaks the decoder-only symmetry every frontier LLM has relied on since the transformer was invented. The result is a model that activates only 8B parameters per input token and 16B per output token, while DeepSeek claims it now beats its own 1.6T flagship V4-Pro on GPQA Diamond, Codeforces, and Terminal-Bench 2.1.
The practical story is even sharper than the benchmark story. KV cache storage for a live session drops to roughly one-quarter of the HBM and one-eighth of the SSD footprint that V4-Flash needed. The original DeepSeek V1's KV cache is now 437× larger than V4.1-Flash's. And because the pricing table was designed alongside the architecture, off-peak rates are exactly half of peak, cached input costs 1/50th of cache miss, and output is the most expensive tier by a wide margin. Four days after launch, DeepSeek routed all deepseek-v4-pro traffic to this model at these prices. If your production code still says deepseek-v4-pro, you are already running V4.1-Flash.
What's New

552B MoE, but asymmetric activation: V4.1-Flash is 552B total parameters — the smallest member of DeepSeek's new architecture family, yet larger than the 284B V4-Flash it replaces. The asymmetry is the point: 8B active per token during prefill (input ingestion) and 16B active during decode (output generation). A conventional decoder-only MoE activates the same expert set for both phases; CED does not.
Native multimodal vision, built in: Vision understanding is no longer a separate branch. The experimental deepseek-v4-flash-vision-exp line that launched August 21 is folded directly into the main V4.1-Flash model — no separate model ID, no bolt-on visual encoder workflow.
MIT-licensed weights, same-day: Open weights went live on Hugging Face the day of release under the MIT license, with the technical report PDF linked from the same repo. The model is callable on the API as deepseek-flash — note the shortened name, not deepseek-v4.1-flash.
V4-Pro already retired in practice: Old model names deepseek-v4-flash and deepseek-v4-flash-vision-exp are now routed to V4.1-Flash. Starting September 14, 2026 at 04:00 UTC, deepseek-v4-pro was routed to V4.1-Flash too, billed at V4.1-Flash rates, until a V4.1-Pro ships.
Architecture

CED is a 40-layer transformer split in half: the first 20 layers form a causal encoder, the remaining 20 form a decoder. The split itself echoes T5-era sequence-to-sequence design. Two choices make it genuinely new.
The encoder is causal, not bidirectional. Classic encoder-decoder designs (BERT, T5) use a bidirectional encoder where every prompt token attends to every other. CED's encoder uses the same left-to-right causal masking a decoder-only model uses. That looks like giving up information — but it is what makes encoder hidden states valid prefixes of each other. When an agent extends a prompt with more retrieved context, the representation built from the first chunk is reused rather than recomputed. A bidirectional encoder invalidates everything it computed the moment a new token arrives; CED does not.
The decoder reads one projected KV, not 20 per-layer caches. A standard 40-layer decoder-only model maintains 40 independent KV caches, one per layer. In CED, the decoder's global KV cache is projected from the encoder's final hidden states — a single representation computed once during prefill. The decoder spends its 20 layers generating the next token well instead of re-deriving "what does this prompt mean" at 20 separate depths.
That is the structural reason the active counts land where they do. Compare:
| Spec | V4.1-Flash (CED) | V4-Pro (prior flagship) |
|---|---|---|
| Total parameters | 552B | ~1.6T |
| Active, prefill (input) | 8B | ~49B (flat) |
| Active, decode (output) | 16B | ~49B (flat) |
| Layers | 40 (20 enc + 20 dec) | decoder-only, per-layer KV |
| Context window | 4K → 1M | 1M |
| KV cache, HBM | 1/4 of V4-Flash | baseline |
| KV cache, SSD | 1/8 of V4-Flash | baseline |
| License | MIT | MIT |
Three deployment tricks stack on top of the architecture. CSA2 (Compressed Sparse Attention 2) reuses the global KV cache and Top-K index across layers, so each layer keeps only its own query vectors and local attention cache. Global KV is stored in FP4 from training, with negligible quality loss. And SWA Bounded Replay reconstructs sliding-window attention state by replaying only the most recent n_win tokens instead of persisting it to SSD — which is the specific mechanism behind the 8× SSD reduction. The combined effect: stretching context from 4K to 1M raises single-token decode FLOPs by only about 25%.
Pricing

The rate card is precise, and the ratios matter more than the absolute numbers:
| Token type | Off-peak | Peak |
|---|---|---|
| Input, cache hit | $0.003 / M | $0.006 / M |
| Input, cache miss | $0.15 / M | $0.30 / M |
| Output | $0.60 / M | $1.20 / M |
Peak hours are narrow and explicit: 01:00–04:00 UTC and 06:00–10:00 UTC, Monday–Friday. Every other hour, plus the entire weekend, is off-peak, and off-peak is exactly half of peak in every tier.
Two ratios drive build decisions. A cache-hit input token costs 1/50th of a cache miss — the cheapest possible way to repeat a stable system prompt, repo snapshot, or long document across turns. Output is the most expensive tier by a wide margin: 4× cache-miss input, 200× cache-hit input. The pricing is not accidental. It rewards exactly the workload CED was built for — shove cost into cheap, cacheable input, keep generated output short.
For an Americas team, those two UTC windows land in late-night and mid-morning US hours, so a large share of business traffic defaults to off-peak. East Asia and European teams are more likely to hit a peak window; schedule batch work accordingly.
Benchmarks

DeepSeek's own changelog reports three headline numbers:
| Benchmark | V4.1-Flash | What it measures |
|---|---|---|
| GPQA Diamond | 90.9 | Graduate-level science reasoning, anti-memorization |
| Codeforces rating | 3471 | Real competitive programming, human rating scale |
| Terminal-Bench 2.1 | 90.6 | Real terminal/command-line agentic tool use |
DeepSeek states these put V4.1-Flash ahead of V4-Pro, GLM5.3, and Kimi-K3 on the Agentic Benchmark across performance, cost, speed, and total runtime. Two caveats are worth holding. These are self-reported numbers — DeepSeek chose the tasks and ran the eval, with no independent third-party replication at launch. And the claim is structurally interesting: a 552B model with 8–16B active per token beating a 1.6T / 49B-active flagship on real reasoning, coding, and terminal agent tasks. If independent labs corroborate that in the coming weeks, it is a real signal about how little active compute you need when prefill and decode stop being the same operation.
What It Means for Developers
Agentic coding and research pipelines are the native workload. A coding agent that ingests 400,000 tokens of repo context and emits a 1,500-token patch has a 267:1 input-to-output ratio. CED spends that massive prefill at 8B active; a flat 49B decoder spends it at 49B. The savings compound across thousands of requests — and the 8× smaller persistent SSD cache means long, paused-and-resumed agent sessions stay affordable over hours and days instead of forcing a full re-prefill.
Cache-hit engineering is now a primary cost lever. With input cache hit at 1/50th of miss, structure prompts so unchanging context (system prompt, stable retrieval set, the same source document) lands as a hit on every call after the first. Combined with off-peak batching, the effective price for repetitive agentic workloads collapses well below the nominal $0.15/M.
Check your model strings today. Any production system hardcoding model="deepseek-v4-pro" is no longer calling V4-Pro — it is calling V4.1-Flash under a legacy alias, likely cheaper, but with a different latency and cost profile than you load-tested. Update explicitly to model="deepseek-flash" and re-baseline.
Self-hosting got more practical. At 8B prefill / 16B decode active parameters and MIT weights, a 552B CED model is dramatically cheaper to serve than its active count suggests. DeepSeek is actively working with the open-source community on inference adaptation and offers direct support for deployments with ~2k-GPU clusters and storage infrastructure. DeepJIT, a new xPU kernel JIT library supporting both NVIDIA CUDA and Huawei Ascend NPU backends, shipped alongside it.
Bottom Line
V4.1-Flash is the first credible break from decoder-only symmetry at the frontier. By making the causal encoder do input understanding once and the decoder read one projected KV, DeepSeek got a 552B model that activates 8B on input and 16B on output, shrank KV cache to 1/4 HBM and 1/8 SSD, and claims flagship-beating reasoning at aggressive peak/off-peak pricing. The quiet retirement of V4-Pro under the same model name is the bigger signal: DeepSeek is betting that CED is the architecture every future model in its line will be built on. Watch for independent benchmark corroboration and the inevitable V4.1-Pro.
Resources
- Official Announcement — DeepSeek V4.1-Flash launch post and architecture rationale
- DeepSeek API Docs — Model changelog, migration, and model-ID routing notes
- API Pricing — Peak/off-peak rate card and UTC peak windows
- Hugging Face — DeepSeek-V4.1-Flash — MIT weights and technical report PDF
DeepSeek-V4.1-Flash is available via API at platform.deepseek.com under the model name deepseek-flash, with open weights on Hugging Face under the MIT license.