DeepSeek Launches Smaller, Faster V4.1-Flash Model That Beats Its Own Flagship
The Chinese AI startup has released DeepSeek-V4.1-Flash, a compact model that outperforms the larger V4-Pro across multiple benchmarks while cutting API costs by roughly 70 percent.

Hangzhou DeepSeek Artificial Intelligence Basic Technology Research Co. Ltd., a Chinese artificial intelligence startup, unveiled DeepSeek-V4.1-Flash on Sept. 14, positioning it as the entry point in a fresh architecture lineup. According to the company, independent testing demonstrates that this open-weight model surpasses its substantially larger sibling, DeepSeek-V4-Pro, in performance metrics, expenses, latency, and overall throughput.
Starting immediately, all API calls directed to V4-Pro will be rerouted to V4.1-Flash and charged at the smaller model's pricing structure, a change that remains in effect until DeepSeek releases a V4.1-Pro variant. The earlier V4-Flash and the experimental vision model introduced in August have both been discontinued; traffic to either now routes to V4.1-Flash instead.
Architecture and Technical Improvements
V4.1-Flash employs a mixture-of-experts architecture containing 552 billion total parameters, roughly double the 284 billion found in V4-Flash. The model's design activates only 8 billion parameters during prompt processing and 16 billion during response generation, thanks to a novel causal encoder-decoder structure. Vision capabilities, previously available only through the experimental August release, are now integrated directly into the base model.
Substantial optimization work centered on reducing the key-value cache footprint. DeepSeek's technical documentation reveals that the model encodes these values using four-bit floating-point representation, achieving a global memory requirement of 890 bytes per token—approximately one-quarter of V4-Flash's consumption. Persistent cache storage on solid-state drives requires roughly one-eighth the space needed by the prior generation.
Performance Benchmarks
DeepSeek's benchmark comparison pits V4.1-Flash at maximum reasoning intensity against Anthropic PBC's Claude Opus 5 and OpenAI Group PBC's GPT-5.6 Sol. On Terminal-Bench 2.1, V4.1-Flash achieved 90.6 points, marginally surpassing Opus 5's 89.1 and GPT-5.6 Sol's 88.8. For the DeepSWE v1.1 software engineering evaluation, V4.1-Flash resolved 74.2% of tasks versus 74% for Opus 5 and 62.7% for V4-Pro. The American models maintain their edge on the GPQA Diamond science reasoning assessment.
Pricing and Availability
Off-peak API rates stand at 15 cents per million uncached input tokens and 60 cents per million output tokens, with pricing doubling during weekday peak hours. Developers currently using V4-Pro face charges of $3.96 per million output tokens during peak times, whereas V4.1-Flash costs $1.20—representing a savings of roughly 70% on output token expenses.
Model weights are accessible via Hugging Face under the MIT license. The model is now operational in DeepSeek's web and mobile applications. DeepSeek indicated plans to collaborate with the open-source community on inference optimization and to investigate additional deployment pathways.
Broader Context
The announcement coincides with Anthropic's release of a threat intelligence report naming DeepSeek among seven China-based AI laboratories that Anthropic claims conducted distillation campaigns targeting Claude. Anthropic documented more than 12.1 million interactions attributed to DeepSeek over a 14-day span in July.
DeepSeek emerged from Chinese hedge fund High-Flyer. Founder Liang Wenfeng reportedly invested $3 billion into a funding round exceeding $7.4 billion in June, which valued the organization above $50 billion.


