DeepSeek recently cut prices for its V4 Flash model and released a refreshed version called V4 Flash 0731. This new model features 284 billion parameters. Simultaneously, OpenAI discounted its GPT-5.6 Luna model by as much as 80 percent to stay competitive.
OpenAI priced GPT-5.6 Luna at $0.2 per 1 million input tokens and $1.20 per 1 million output tokens. DeepSeek’s V4 Flash 0731 is priced lower at $0.14 per 1 million input tokens and $0.28 per 1 million output tokens, supported by 20,000 NVIDIA H100 GPUs.
DeepSeek informed customers that API prices will soon undergo a significant increase. This change specifically affects API users rather than those running open-weight models on their own infrastructure, marking a shift in the company's pricing strategy.