본문으로 건너뛰기
All news

DeepSeek V4-Flash-0731 Official Release: Beats Its Own Pro Model on 9 Agent Benchmarks

Summary: DeepSeek moved V4-Flash from preview to official release on July 31, shipping under the identifier DeepSeek-V4-Flash-0731. The same 284B/13B MoE architecture was re-post-trained, and the result beats the company's own Pro-Preview on every agent and coding benchmark published.

Key Facts

  • Benchmark sweep: Terminal Bench 2.1 score of 82.7 vs. Pro-Preview's 72.1. DeepSWE jumped from 7.3 to 54.4; DSBench-FullStack from 37.0 to 68.7.
  • Zero migration cost: The deepseek-v4-flash API endpoint is unchanged — existing integrations get the upgrade automatically.
  • Pricing: $0.14/M input tokens (cache miss), $0.003/M on cache hits, $0.28/M output — unusually cheap for frontier-class agent performance.
  • Caveat: As of July 31, no independent lab had reproduced these figures; all numbers are vendor-reported.

Why It Matters

When a Flash-tier model outperforms its own Pro flagship on agentic tasks, the cost-performance calculus for large-scale deployments shifts again. DeepSeek is making it increasingly hard to justify frontier-model pricing for coding and agent workloads — a pattern that is forcing price cuts across the industry.

Read More

뉴스레터 구독

무료 뉴스레터

매주 핵심 AI 소식, 한 번에 받기

쏟아지는 AI·LLM 뉴스 중 꼭 알아야 할 것만 골라 메일로 보내드려요. 뉴스레터 발송이 시작되면 구독자분들께 가장 먼저 보내드립니다.