DeepSeek launched the V4-Flash 0731 public beta API on July 31, with official claims that its upgraded agentic capabilities have surpassed V4-Pro-Preview, and it supports the Responses API format, fully compatible with Codex. However, the company clarified that improvements are limited to the Flash API; V4-Pro API/App/Web remain unchanged, and the full V4-Pro release is still pending.
Community observers quickly quantified the leap: @cline noted a Terminal-Bench score of 82.7, up 25.8 points from the April preview's 56.9. Artificial Analysis reported that the model's index rose from 40 to 50, just 1 point behind GPT-5.6 Luna (max 51), while cost per task is about 60% lower. Other agentic benchmarks also improved significantly: GDPval-AA v2 Elo from 1189 to 1559, Terminal-Bench 2.1 at 79%, τ³-Bench Banking up 8 percentage points, and output token usage down 12%.
Notably, these gains are claimed to be achieved without changing architecture or scale. Artificial Analysis confirmed that V4-Flash 0731 remains 284B total params/13B active, 1M context, text-only, priced at $0.14/$0.28 per million input/output tokens, with an unusually aggressive 98% cache hit discount, bringing cached token price to $0.0028 per million.