Lifestyle

DeepSeek-V4.1-Flash Launches, Mixed Word on V4-Pro: What to Watch When Models Switch

DeepSeek launched V4.1-Flash on September 10, 2026, temporarily routing the old V4-Flash model names to it. The release announcement says V4-Pro would be rerouted from September 14; the change log and pricing page say it stays available after September 14. Using official pages, this article sets out the order of the two statements and the status when checked, and explains how people using DeepSeek indirectly through third-party tools can spot a model switch and read pricing and privacy terms.

Updated: About 9 min read

Original illustration of a model name that stays the same while the model behind it has already been swapped
Image: Mokaair (© Mokaair)

On September 10, 2026, DeepSeek published DeepSeek-V4.1-Flash in its official API documentation, describing it as the smallest model in its new architecture family, with native image understanding; it is called with the model name deepseek-flash. The previous-generation V4-Flash and V4-Flash-Vision-Exp were retired at the same time, and their old names are temporarily routed to the new model. The announcement also said that requests to V4-Pro would be rerouted to V4.1-Flash starting September 14, but at the time of checking, DeepSeek's official change log and pricing page said V4-Pro would continue to be offered after September 14.

This article was checked on September 15, 2026. The announcement describes changes to the API; these official pages do not say which model the DeepSeek app and web version currently use, and this article makes no assumption about it. This site has not called or benchmarked the model, so the capability and benchmark claims here are DeepSeek's own statements, and the everyday and work scenarios are examples designed by the editors.

What the V4.1-Flash Announcement Says

DeepSeek says V4.1-Flash is a mixture-of-experts (MoE) model with 552B parameters built on a new Causal Encoder-Decoder architecture. It activates 8B parameters when reading input and 16B when generating output, which the company describes as asymmetric between input and output. The announcement also says a new pre-training approach, combined with larger-scale reinforcement learning post-training, lets it beat flagship models including DeepSeek-V4-Pro on benchmarks. That is the vendor's claim: the comparison table in the model card also shows items where it trails models from other companies, and this site has not rerun any tests.

The other headline is caching. When a model handles a long conversation, it temporarily stores intermediate computation data, known as the KV cache. The announcement says that compared with the previous generation, V4.1-Flash cuts the KV cache's HBM requirement to one quarter and its SSD requirement to one eighth, and that the more efficient architecture is the reason API prices have been lowered.

On the weights, the announcement links to the official Hugging Face page. The model card states that the repository and model weights are licensed under the MIT License, that the model supports a context of up to one million tokens, and that it can take images and text directly and output text. Being downloadable does not mean an ordinary computer can run it: the inference example linked from the model card runs with an 8-way model-parallel setup. The release announcement separately invites organizations planning large-scale deployments with 2,000 GPUs plus a storage cluster to contact DeepSeek.

Old Names Rerouted and the Future of V4-Pro: Two Statements in Sequence

The first statement comes from the September 10 release announcement. It says multiple tests showed V4.1-Flash beating V4-Pro on performance, cost, speed, and total time taken, so DeepSeek planned to retire V4-Pro: starting at UTC 04:00 on September 14, 2026 (12 noon Taiwan time), all deepseek-v4-pro requests would be handled by V4.1-Flash and billed at V4.1-Flash prices until V4.1-Pro is released.

The second statement appears in the change log and on the pricing page. On the day of checking, the September 10 entry in the change log and a note on the pricing page said that in response to user demand, DeepSeek had decided to keep offering the V4-Pro API service after September 14, 2026, with billing unchanged, and that any changes would be announced separately. On the pricing page, the version listed for deepseek-v4-pro is DeepSeek-V4-Pro-0813. The wording “decided to keep offering” suggests a decision made after the rerouting plan, but the pages carry no date, and the news list in the API documentation has no separate announcement.

When checked on September 15, both the English and Simplified Chinese versions of the release announcement still carried the original text about rerouting starting September 14. Someone who reads only the announcement may think V4-Pro has already been replaced. According to the change log and pricing page, deepseek-v4-pro continues to be offered after September 14, while the two old V4-Flash names are served by V4.1-Flash and billed at Flash prices.

Checked on September 15, 2026; compiled from DeepSeek's change log and models and pricing page, official API only
Model nameModel actually servingPricing
deepseek-flashDeepSeek-V4.1-FlashFlash pricing
deepseek-v4-flashV4.1-Flash (original model retired)Flash pricing
deepseek-v4-flash-vision-expV4.1-Flash (original model retired)Flash pricing
deepseek-v4-proDeepSeek-V4-Pro-0813Billing unchanged

Why People Who Never Call the API Directly Are Affected Too

Most readers in Taiwan will never call the DeepSeek API themselves, but they may still be using it indirectly. Some translation extensions, writing apps, and internal company Q&A tools put a DeepSeek model name in their settings and forward users' input to it; DeepSeek's announcement also names WorkBuddy (including CodeBuddy) and OpenCode as fully supporting V4.1-Flash. When a provider routes an old name to a new model, nothing in the tool's settings has to change, yet the model answering behind the scenes is already a different one.

This release is a clear example: the name deepseek-v4-flash says V4, but the model actually serving it is V4.1-Flash. The change log also records that when V4 launched on April 24, the two old names deepseek-chat and deepseek-reasoner pointed during the transition to the non-thinking and thinking modes of deepseek-v4-flash, respectively, and that they were scheduled to be discontinued on July 24, 2026.

Take a scenario designed by the editors: an office worker uses an internal company summarizer every day to tidy up meeting notes, and after September 10 notices that the summaries are longer and broken into sections differently. A good first step is to check the tool's update notes or ask an administrator whether the backend connects to DeepSeek and which model name is configured. A sudden change in answer style, length, output format, or image-reading ability is a signal worth asking about. It does not necessarily mean things got worse, but if the output goes into formal documents, it is wise to spot-check samples again.

Four checks when a model changes hands: read the tool's update notes, compare answers, confirm the model name and version, read the privacy policy
Four checks when a model changes hands: read the tool's update notes, compare answer length and format, confirm the model name and actual version, and read the official privacy policy. · Image: Mokaair (© Mokaair)

Pricing, Peak Hours, and What a Fixed Version Means

API pricing gets only a brief mention here. According to the pricing page, during peak hours deepseek-flash costs US$0.3 per million input tokens (cache miss) and US$1.2 per million output tokens; for deepseek-v4-pro the figures are US$1.32 and US$3.96. Off-peak prices are half the peak rate, and peak hours run from UTC 01:00 to 04:00 and from 06:00 to 10:00, Monday through Friday, which is 9 a.m. to 12 noon and 2 p.m. to 6 p.m. Taiwan time. The announcement says the new prices took effect at UTC 04:00 on September 10.

When an old name is routed to a new model, the bill changes as well. The pricing page says requests using the old V4-Flash names are billed at Flash prices. Had V4-Pro been rerouted as originally planned, the pricing page figures mean people using V4-Pro would have paid less, but their answers would have come from a different model. Features differ too: the pricing page marks deepseek-flash as supporting image understanding and deepseek-v4-pro as not supporting it. For people paying a monthly fee for a third-party tool, whether the plan price changes as a result depends on each tool's own documentation.

A fixed version means a tool specifies a model version that will not be quietly swapped out, which makes its behavior more predictable. DeepSeek's pricing page lists model names and actual versions separately, but does not say whether a dated version name can be used in calls; the August 13 change log entry also says that sticking with the name deepseek-v4-pro gets you the latest version. People who manage their own tools can record the model name, version, and date checked, and review the change log regularly.

Terms to Read Yourself Before Using a Chinese Company's Service

When you use DeepSeek through its official API or app, what you type is sent to DeepSeek's service for processing. For how data is stored, where it is kept, and whether it is used to improve the service, read DeepSeek's official privacy policy and terms of service; this article does not interpret them for you. Before a company sends internal data to any outside AI service, whatever country the provider is based in, it should first confirm that its own data policies allow it.

Indirect use calls for one more layer of questions. A third-party tool may connect directly to DeepSeek's official API, or it may use a model hosted on another cloud platform, and different terms apply to each. Look in the tool's help pages or privacy policy for sections on things like “model providers” or “data transfer.” Until you find an answer, do not paste in national ID numbers, customer lists, or unpublished contracts.

Open weights offer another option: the model card uses the MIT License, so organizations can deploy the model in their own environment without sending data to DeepSeek's API. Self-hosting, however, means taking on the hardware, operations, and data management, trade-offs this site has already covered in its article on Qwen3.5 open weights.

Latest travel guides

Sources

Lifestyle