DeepSeek V4-Flash-0731 Official Release: Beats Its Own Pro Model on 9 Agent Benchmarks
Summary: DeepSeek moved V4-Flash from preview to official release on July 31, shipping under the identifier DeepSeek-V4-Flash-0731. The same 284B/13B MoE architecture was re-post-trained, and the result beats the company's own Pro-Preview on every agent and coding benchmark published.
Key Facts
- Benchmark sweep: Terminal Bench 2.1 score of 82.7 vs. Pro-Preview's 72.1. DeepSWE jumped from 7.3 to 54.4; DSBench-FullStack from 37.0 to 68.7.
- Zero migration cost: The
deepseek-v4-flashAPI endpoint is unchanged — existing integrations get the upgrade automatically. - Pricing: $0.14/M input tokens (cache miss), $0.003/M on cache hits, $0.28/M output — unusually cheap for frontier-class agent performance.
- Caveat: As of July 31, no independent lab had reproduced these figures; all numbers are vendor-reported.
Why It Matters
When a Flash-tier model outperforms its own Pro flagship on agentic tasks, the cost-performance calculus for large-scale deployments shifts again. DeepSeek is making it increasingly hard to justify frontier-model pricing for coding and agent workloads — a pattern that is forcing price cuts across the industry.
Read More
- TechTimes Benchmark Analysis — TechTimes
- Benchmarks & Pricing Summary — CosmicJS