
DeepSeek V4-Flash Is Now the Cheapest AI Model to Run, Research Firm Says
TLDR
- DeepSeek launched V4-Flash on 1 August 2026, and within days a research firm ranked it the cheapest well-known AI model to run.
- Artificial Analysis says V4-Flash costs roughly 3 US cents (about 15 sen) to complete its full benchmark suite, undercutting every major Western rival.
- Compared with Anthropic’s Claude Opus line, the model is over 100 times cheaper per intelligence-index task, with OpenAI and Google Gemini models priced well above it.
- The release lands in the same week Alibaba unveiled its 2.4-trillion-parameter Qwen3.8-Max, marking a coordinated Chinese push on frontier AI.
- For Malaysian builders, the pricing resets what is realistic for self-hosted and API-driven AI products in 2026.

Chinese AI startup DeepSeek has thrown down a fresh challenge to OpenAI, Anthropic and Google with its new V4-Flash model. Released on 1 August 2026, V4-Flash is the cheapest well-known large language model to run, according to independent benchmarking firm Artificial Analysis.
Artificial Analysis measured V4-Flash at roughly 3 US cents, or about 15 Malaysian sen, to complete its full intelligence-index benchmark suite, which covers coding, reasoning, agentic work and knowledge tasks end to end. Anthropic’s top-tier Claude Opus and Google’s Gemini Ultra land closer to USD 3 to USD 5 per benchmark run, with OpenAI’s GPT-5.6 in between. V4-Flash is therefore over 100 times cheaper than Anthropic on a per-task basis.
Reuters first reported the Artificial Analysis figures on 3 August. DeepSeek has historically positioned itself on price efficiency rather than absolute benchmark wins, and V4-Flash continues that strategy. The release lands in the middle of a wider Chinese AI surge: the same week Alibaba’s Qwen team formally unveiled Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model the company calls its most capable to date, while MiniMax released open weights for its H3 video model.
What this means for Malaysian builders
For Malaysian startups, agencies and solo developers, V4-Flash changes the math on what is buildable. A chat assistant processing 10,000 customer queries a day on a Western frontier model would run into four-figure USD bills monthly; the same workload on V4-Flash via DeepSeek’s API sits in the tens of dollars. That gap matters in a market where ringgit-denominated margins are thin and cloud credits rarely absorb production-scale AI traffic.
There are caveats. Artificial Analysis’s pricing assumes reasonable traffic patterns and cached prompts; very long contexts or unusual workloads can swing costs back toward competitors. DeepSeek’s API is hosted in mainland China, which raises data-residency questions for any Malaysian business handling personal data under PDPA. For local-facing prototypes, marketing tools and internal copilots though, V4-Flash is the cheapest credible large model available right now, and that pressure is likely to push Western providers to cut their own prices within weeks.
Our Take
It is easy to read V4-Flash as another round in the US-China AI race, and partly it is. The more practical story is what 3 cents per benchmark run does to product economics. For a year now, Western frontier model pricing has been the bottleneck for serious AI deployment outside Silicon Valley; every Malaysian team we have heard from has had to ration usage or pre-compute answers offline. A model at this price removes that constraint, even if capability is a tier below GPT-5.6 or Claude Opus.
V4-Flash will not displace frontier models for the hardest reasoning or coding tasks. Artificial Analysis still scores it below DeepSeek’s own V4 and well under Anthropic’s Opus on raw intelligence. What it does is move the floor. Every chatbot, summariser and retrieval workflow that does not need top-tier reasoning now has a credible 100-times-cheaper option, and the interesting question is how fast OpenAI, Google and Anthropic respond, and whether the price drop spreads from the API layer down to enterprise contracts.






