DeepSeek V4 Flash 0731 Explained 2026: $0.14/$0.28, 1M Context, 82.7% Terminal-Bench
DeepSeek shipped DeepSeek-V4-Flash-0731 on July 31: same architecture, but re-post-training pushes coding and agent performance up. An open-weight (MIT) sparse MoE with 13B active/284B total, 1M context, $0.14/$0.28 pricing and 82.7% on Terminal-Bench.
Another open-weight release out of China. On July 31 DeepSeek published DeepSeek-V4-Flash-0731 — weights on Hugging Face and the API in public beta. The point is simple: the architecture is unchanged; re-post-training is what lifts coding and agent performance. Pricing stays aggressive. It’s a sparse MoE with 13B active (284B total), a 1M context window, and $0.14/$0.28 pricing.
What actually changed?
This isn’t a new design — it’s the same model re-trained to perform better. The model card is explicit: architecture and size are unchanged, and the gains come from re-post-training.
- Architecture: sparse MoE, 13B active of 284B total parameters
- Context: 1M tokens, up to 65,536 output tokens
- Weights released under the MIT license — free for commercial use
- Tool use, function calling, vision input and step-by-step reasoning in one model

The push is to run coding, reasoning and agents from a single model instead of stitching services together. For the broader China open-model wave, see China AI via OpenRouter rankings and Kimi K3.
How cheap is it?
Per million tokens — an order of magnitude below the top frontier models.
| Item | DeepSeek V4 Flash 0731 |
|---|---|
| Input | $0.14 |
| Output | $0.28 |
| Cache input | $0.003 |
The heavier your agent or coding calls, the wider the real-bill gap. On price alone it stands in the same ring as low-cost coders like Grok 4.5.

Where is it strong?
It stands out on practical, hands-on benchmarks more than on raw intelligence.
- Artificial Analysis Intelligence Index 50 (well above the ~25 median for similar open-weight models)
- Terminal-Bench 2.1 82.7%
- Toolathlon 70.3
Terminal, agent and coding scores are the strength, and it beats its own larger Pro model on some agent benchmarks.

Frequently Asked Questions
Q. Is it a brand-new model? No. Architecture and size match the previous version; the gains come from re-post-training. It’s the official release superseding the preview.
Q. Can I use it commercially? Yes — the weights are MIT-licensed, so commercial use is free. You can pull them from Hugging Face, and a public beta API is open.
Q. What is it good for? Coding, reasoning and agent workflows. It combines tool use, function calling, vision and reasoning in one model, suiting teams that want to consolidate instead of chaining services.
Related posts
Big TechOpenAI's Astra Explained: Multi-Agent Model Family That Cracked 10 Open Math Problems
From answer engine to research partner — a model built to stay on one problem for days
Big TechAmazon Soars 13% on AWS's 37% Growth While Apple Sinks 7%: Earnings Week Wrap
Earnings week's final verdict: show AI revenue and surge, or get caught by AI's side effects and sink
Big TechMicrosoft vs Meta Q2 2026 Earnings: Azure +43% While Meta's Profit Falls 14%
Same AI spending story, opposite stock moves: +10% vs -10% — Wall Street wants receipts