DeepSeek's V4.1-Flash Cuts Memory and API Costs While Claiming Performance Edge Over V4 Pro
DeepSeek has unveiled V4.1-Flash, a streamlined model that the company says reduces memory demands and API expenses while surpassing its predecessor on internal benchmarks. The move includes an aggressive shift that will route V4 Pro requests to the new model at lower pricing.

On Thursday, Chinese AI firm DeepSeek introduced V4.1-Flash as the entry-level offering within its newly unveiled V4.1 architecture lineup. The model incorporates built-in visual capabilities alongside architectural improvements aimed at boosting performance speed, request throughput and operational efficiency.
The system contains 552 billion parameters structured as a mixture-of-experts configuration, though DeepSeek reports that roughly 8 billion parameters activate per input token and 16 billion per output token. The company credits its novel Causal Encoder-Decoder design, enhanced training methods and expanded reinforcement learning scale for enabling strong performance without requiring the entire model to process each request.
V4.1-Flash supports a context window of one million tokens, making cache efficiency a critical factor for extended dialogues and autonomous agent applications.
Benchmark Performance
According to DeepSeek's internal testing, V4.1-Flash exceeded V4 Pro and rival systems across select coding, security and agent benchmarks. On Terminal-Bench 2.1, the model achieved 90.6 points, compared to 88.8 for OpenAI's GPT-5.6 Sol, 88.3 for Moonshot AI's Kimi K3 and 87.9 for V4 Pro.
The model also registered 88.1 on Cybergym, surpassing V4 Pro at 83.3, Kimi K3 at 80 and GPT-5.6 Sol at 84.5. On DeepSWE v1.1, V4.1-Flash narrowly outpaced Anthropic's Claude Opus 5, though that Anthropic model maintained advantages in other assessments.
These findings stem from DeepSeek's own evaluation and may not represent independently verified outcomes or actual production scenarios. DeepSeek has made model weights available via Hugging Face under an MIT license, enabling developers to conduct independent assessments subject to repository guidelines.
The Real Change is in Memory
DeepSeek's most compelling advantage may lie beyond benchmark comparisons. The firm claims V4.1-Flash requires just one-quarter of the HBM and one-eighth of the SSD storage needed for the prior generation's key-value cache. According to SCMP, the cache footprint decreased from 3,514 bytes per token in the earlier Flash iteration to 890 bytes.
This distinction carries weight for AI agents, where preserving extensive context can represent a substantial cost when processing millions of queries. Reduced memory consumption could enable organizations to process more simultaneous workloads without expanding infrastructure.
DeepSeek has also reduced API rates, with off-peak cached input priced as low as 0.02 yuan per million tokens. Final expenses will fluctuate based on usage patterns and the proportion of cached input, uncached input and output tokens.
DeepSeek is retiring its own flagship
The company is executing a notably bold restructuring of its current product portfolio. Beginning September 14, DeepSeek will redirect V4 Pro requests to V4.1-Flash and apply Flash pricing until V4.1-Pro becomes available.
Given that the DeepSeek V4.1 Flash model comprehensively surpasses the V4 Pro in all metrics … it would not be appropriate to provide DeepSeek users with the originally underperforming V4 Pro model at a higher price
Cui Tianyi, head of DeepSeek's Harness team, according to SCMP
Older V4-Flash and V4-Flash-Vision-Exp endpoints will also be discontinued and temporarily redirected to V4.1-Flash.
Teams relying on these endpoints should validate V4.1-Flash functionality ahead of the routing transition, particularly if their systems depend on specific output structures, latency thresholds or certified model versions.
What this means for AI buyers
DeepSeek is wagering that AI purchasers prioritize practical value per dollar and execution speed over raw model dimensions. This approach creates competitive pressure on rivals to enhance both model capabilities and the financial efficiency of large-scale deployment.
For enterprises developing code assistants, retrieval systems or autonomous agents, a model maintaining extended contexts while consuming substantially less memory might prove more valuable than a larger competitor that excels in isolated tests.
V4.1-Flash's reduced expenses and memory footprint could prove appealing for large-scale coding, search and agent operations, though DeepSeek's stated improvements may differ in practical implementations. Organizations evaluating the model or impacted by the V4 Pro routing change should assess output fidelity, response times, system compatibility and total infrastructure spending relative to current solutions before transitioning.


