DeepSeek V4 Pro ranks 12th on Vals Index, excels in cybersecurity
Synopsis
Key Takeaways
DeepSeek, the Hangzhou-based artificial intelligence start-up, quietly released DeepSeek-V4-Pro-0813 on Wednesday, 13 August 2026, billing it as a flagship update with 'significantly enhanced agent capabilities' — but early benchmark results show the model struggling against top-tier rivals, even as it earns praise from cybersecurity researchers.
A stealth launch with a vanishing statement
The update arrived without a formal announcement, accompanied only by a brief note on DeepSeek's official website that was removed by Thursday afternoon. The release follows last month's DeepSeek V4 Flash, a smaller model that rattled Silicon Valley with its extreme cost efficiency. Developers, however, have been less enthusiastic about the new flagship, citing underwhelming overall performance and disappointing pricing.
Benchmark struggles against rivals
DeepSeek-V4-Pro-0813 scored 53 on the Artificial Analysis Intelligence Index, matching Zhipu AI's GLM-5.2 released in June but falling four points short of the mid-tier Terra model in OpenAI's latest GPT-5.6 series and seven points behind Moonshot AI's Kimi K3. On the Vals Index — compiled by San Francisco-based Vals AI to evaluate models across multiple benchmarks — the model ranked 12th, trailing OpenAI's previous-generation GPT-5.5 and lagging well behind frontier systems including Kimi K3 and Anthropic's Claude Opus 5.
According to Vals AI, the model struggled most in two specific areas: completing tasks within a sandboxed terminal environment and generating complex financial models in Excel spreadsheets. These weaknesses are notable given that agentic task execution is precisely the capability DeepSeek highlighted in its now-deleted launch statement.
Why it matters: cybersecurity as a bright spot
Despite the mixed overall showing, researchers in niche domains — particularly cybersecurity — have reportedly found DeepSeek-V4-Pro-0813 impressive. This mirrors a pattern seen with earlier DeepSeek releases, where domain-specific strength compensates for broader benchmark shortfalls. The cybersecurity performance could make the model attractive to enterprise security teams even if it fails to displace general-purpose leaders.
The competitive backdrop
DeepSeek's flagship now competes in a crowded field where OpenAI, Anthropic, and domestic Chinese rivals like Moonshot AI and Zhipu AI are all pushing capability boundaries. The V4 Flash model demonstrated that DeepSeek can disrupt on cost; the V4 Pro update suggests the company has not yet matched that disruption at the capability frontier. Industry analysts noted that a rank of 12th on the Vals Index represents a significant gap from the podium positions held by Kimi K3 and Claude Opus 5.
What's next
The removal of DeepSeek's launch statement within 24 hours raises questions about whether the company considers this release final or a staged rollout pending further tuning. Developers and enterprise buyers will be watching closely for a revised pricing structure and whether subsequent updates can close the gap on sandboxed agentic tasks — the benchmark category most directly tied to real-world deployment value.