DeepSeek Launches V4.1-Flash With Lower Memory and API Costs

DeepSeek Launches V4.1-Flash With Lower Memory and API Costs

DeepSeek launches V4.1-Flash. Image: Solen Feyissa/Unsplash

DeepSeek V4.1-Flash promises lower memory use and API costs, but buyers should test its performance, compatibility and total deployment expenses.

Sep 11, 2026

DeepSeek says its latest model lowers memory requirements and API costs while outperforming V4 Pro on several internal benchmarks.

Chinese AI company DeepSeek launched V4.1-Flash on Thursday as the smallest model in its new V4.1 architecture family, combining native visual understanding with an architecture designed to improve speed, throughput and serving costs.

The model has 552 billion total parameters in a mixture-of-experts (MoE) system, but DeepSeek says it activates approximately 8 billion parameters per input token and 16 billion per output token. DeepSeek says its new Causal Encoder-Decoder architecture, combined with new pretraining methods and larger-scale reinforcement learning, allows the model to deliver stronger results without using the full model for every request.

DeepSeek also gives V4.1-Flash a one-million-token context window, making its cache efficiency particularly important for long-running conversations and agent workloads.

In benchmark results published by DeepSeek, V4.1-Flash outperformed V4 Pro and several competing models on selected coding, cybersecurity and agent evaluations. On Terminal-Bench 2.1, it scored 90.6, compared with 88.8 for OpenAI’s GPT-5.6 Sol, 88.3 for Moonshot AI’s Kimi K3 and 87.9 for V4 Pro.

The model also scored 88.1 on Cybergym, ahead of V4 Pro at 83.3, Kimi K3 at 80 and GPT-5.6 Sol at 84.5. On DeepSWE v1.1, V4.1-Flash narrowly beat Anthropic’s Claude Opus 5, although the Anthropic model remained ahead on other evaluations.

Those results come from DeepSeek’s own testing and may not reflect performance under independently reproduced conditions or real-world workloads. DeepSeek has released the model weights on Hugging Face under an MIT license, allowing developers to conduct their own evaluations subject to the repository’s published terms.

The real change is in memory

DeepSeek’s bigger selling point may be what happens outside the benchmark chart. The company says V4.1-Flash uses only one-quarter of the HBM and one-eighth of the SSD storage required for its previous generation’s key-value cache. SCMP reports that the cache footprint fell from 3,514 bytes per token in the previous Flash model to 890 bytes.

Advertisement

That matters for AI agents, where maintaining large amounts of context can become an expensive part of running millions of requests. Lower memory requirements could let businesses handle more concurrent workloads without simply throwing more hardware at the problem.

DeepSeek has also cut API pricing, saying off-peak cached input can cost as little as 0.02 yuan per million tokens. Actual costs will depend on usage time and the mix of cached input, uncached input and output tokens.

More must-read AI coverage

DeepSeek is retiring its own flagship

The company is making an unusually aggressive move with its existing lineup. Starting Sept. 14, DeepSeek says V4 Pro requests will be routed to V4.1-Flash and charged at the Flash rate until V4.1-Pro launches.

“Given that the DeepSeek V4.1 Flash model comprehensively surpasses the V4 Pro in all metrics … it would not be appropriate to provide DeepSeek users with the originally underperforming V4 Pro model at a higher price,” Cui Tianyi, head of DeepSeek’s Harness team, said on X, according to SCMP.

Older V4-Flash and V4-Flash-Vision-Exp endpoints are also being retired and temporarily routed to V4.1-Flash.

Organizations using the affected endpoints should test V4.1-Flash before the routing change, particularly if their applications depend on consistent output formats, latency targets or an approved model version.

What this means for AI buyers

DeepSeek is betting that AI customers care less about the size of a model than how much useful work they get for each dollar and how quickly they get it.

That puts pressure on competitors to improve not just model intelligence but the economics of running those models at scale. For companies building coding tools, search interfaces or autonomous agents, a model that can maintain long contexts while using substantially less memory could be more valuable than a larger model that wins isolated benchmarks.

V4.1-Flash’s lower pricing and cache requirements could make it attractive for high-volume coding, search and agent workloads, but DeepSeek’s reported gains may vary across real deployments. Organizations considering the model — or affected by the V4 Pro routing change — should compare its output quality, latency, compatibility and total infrastructure costs against their current deployments before switching.

Advertisement

Read more: US companies testing DeepSeek are finding that lower model prices do not always translate into lower total AI costs.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. He has written for a wide range of technical and business audiences, from IT professionals and cybersecurity leaders to small business owners, executives, and technology buyers. His work has appeared in publications including: TechRepublic eWEEK Channel Insider Geekflare Enterprise Networking Planet eSecurity Planet CIO Insight Webopedia With a background in computer science, Aminu specializes in translating complex technical subjects into clear, practical, and accessible content. His writing helps readers understand emerging technologies, evaluate business software, strengthen cybersecurity strategies, and make more informed decisions about technology investments. Across his work, Aminu focuses on the real-world impact of technology, connecting technical innovation with business value, operational efficiency, security, and long-term digital transformation.