Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

DeepSeek pushes V4-Flash API into public beta at cut-rate pricing

DeepSeek released its V4-Flash API into public beta on July 31, 2026, scoring 82.7 on Terminal-Bench 2.1, near Anthropic's Claude Opus 4.8.

D
Jul 31, 2026 · 1 min read

DeepSeek released its V4-Flash API into public beta on July 31, 2026, a cheaper, agent-tuned model the company says outscores its own larger V4-Pro-Preview on a leading coding benchmark.

On Terminal-Bench 2.1, a test of AI agents working in a command line, DeepSeek-V4-Flash scored 82.7, versus 72.1 for V4-Pro-Preview and 85.0 for Anthropic’s Claude Opus 4.8, according to figures on the model’s Hugging Face card. The scores are DeepSeek’s own and have not been independently verified.

V4-Flash is a Mixture-of-Experts model, an architecture that activates only part of the network per query, with roughly 13 billion active parameters, a 1-million-token context window and re-training for agent and coding work. API pricing is $0.14 per million input tokens and $0.28 per million output tokens, with cached input from $0.0028, undercutting most frontier rivals.

Other reported results show sharp movement on agent tasks: the DeepSWE score jumped to 54.4 from 7.3 in the July preview. The weights ship under an MIT license as DeepSeek-V4-Flash-0731; the V4-Pro API and DeepSeek’s consumer app and web models are unchanged.

The framing to watch is price versus parity. A smaller model beating a larger sibling on benchmarks says as much about post-training as raw scale, and benchmark leads rarely survive contact with real workloads. Still, at roughly a fifth of typical frontier output pricing, V4-Flash sharpens the cost pressure Chinese labs keep applying to U.S. model makers.

More news