DeepSeek Ships V4-Flash-0731, Closing In on Claude Opus 4.8 at 28 Cents Per Million Output Tokens
DeepSeek moved V4-Flash-0731 out of preview on July 31 with sharply higher agentic and coding scores, while holding output pricing at $0.28 per million tokens, well below OpenAI's freshly discounted GPT-5.6 Luna.
DeepSeek has replaced the April preview of V4-Flash with a production release, DeepSeek-V4-Flash-0731, retrained with a post-training pipeline focused on coding, tool use and agentic reasoning. The model keeps its 284B-total, 13B-active mixture-of-experts design and 1M-token context, but its benchmark scores jumped sharply over the preview.
- Terminal Bench 2.1: 82.7, up from 61.8 in the preview, versus Claude Opus 4.8's 85.0
- DeepSWE agentic coding: 54.4, up from 7.3 in the preview
- Artificial Analysis Intelligence Index: 50, ranking #3 of 101 tracked models, well above the 25 median
- Pricing unchanged at $0.14 per million input tokens, $0.0028 on cache hits, $0.28 per million output tokens
OfficeChai notes Opus 4.8 "still leads on every benchmark DeepSeek has published," but the margin has narrowed. On price the comparison isn't close: OpenAI's just-discounted GPT-5.6 Luna runs $0.20/$1.20 per million tokens, still several times V4-Flash-0731's rate. Compare the two directly in coding.