Chips & compute
DeepSeek's V4-Flash is the cheapest major AI model to run
A new model from the Chinese AI lab costs a fraction of what GPT and Claude models cost for similar work, according to research firm Artificial Analysis.
The answer
DeepSeek released V4-Flash, the cheapest well-known AI model to run, on 31 July.
DeepSeek, the Chinese AI lab, has released a new model called V4-Flash that costs far less to run than any other well-known AI system, according to new research.
What happened
DeepSeek officially launched V4-Flash on Friday 31 July. Days later, on 3 August, San Francisco research firm Artificial Analysis reported that the model is by far the cheapest of the well-known AI models to run, according to Reuters.
On paper, V4-Flash charges $0.14 per million "input tokens" (roughly, words you feed into the model) and $0.28 per million "output tokens" (words it generates back). But Artificial Analysis says list prices can be misleading, because a cheap-looking model can still cost more overall if it needs many extra steps to reach the right answer. So it instead measured the average cost of running each model through a standard set of benchmark tests.
By that measure, V4-Flash cost just 3 cents per test. Moonshot's Kimi K3 cost 86 cents, OpenAI's GPT-5.6 Sol cost $1.86, and Anthropic's Claude Fable 5 cost $3.15 — making V4-Flash more than 100 times cheaper to run than Fable 5.
That low cost comes with a trade-off in capability. On the Artificial Analysis Intelligence Index, which combines nine benchmarks covering coding, reasoning and workplace tasks, V4-Flash scored 50 out of 100. That puts it level with Google's Gemini 3.6 Flash, one point behind Meta's Muse Spark 1.1 and Z.ai's GLM-5.2, and at least nine points behind the leading models, Claude Opus 5, Fable 5 and GPT-5.6.
Under the hood, V4-Flash uses a "mixture-of-experts" design, meaning it has about 284 billion total parameters (the internal settings that store what a model has learned) but only activates about 13 billion of them for any single request. It can also handle up to 1 million tokens of context at once, meaning it can process very long documents or conversations in one go, according to The Next Web.
What it means for you
If you use AI tools through an app or business service rather than paying per token yourself, this doesn't change anything for you directly yet. But it matters for the wider market: for routine, high-volume tasks like customer support replies or basic coding help, cost is becoming a bigger factor in which model companies choose, and Chinese labs keep driving that cost down. If services you use start building on cheaper models like V4-Flash, that could eventually mean lower prices, though with somewhat less capable answers than top-tier models provide.
What happens next
DeepSeek is reportedly preparing for a possible stock market listing. The company's earlier R1 model triggered a global sell-off in tech stocks in early 2025, so its moves continue to draw close attention from investors and rival AI labs alike.
Sources
- DeepSeek's new AI model is by far the cheapest of well-known models to run, research firm says — Reuters via Yahoo Finance, 3 August 2026
- DeepSeek's V4-Flash is the cheapest well-known AI model to run, research firm finds — The Next Web, 3 August 2026