DeepSeek V4 Pro Finds 28 of 32 CVEs as Open Models Beat Public Frontiers
Updated
Updated · aikido.dev · Aug 21
DeepSeek V4 Pro Finds 28 of 32 CVEs as Open Models Beat Public Frontiers
2 articles · Updated · aikido.dev · Aug 21
Summary
A 10-model benchmark across 32 fresh vulnerabilities found DeepSeek V4 Pro 0813 delivered the top pooled recall, rediscovering 28 CVEs across three runs after finding 17 on its first pass.
The gain came from repetition: researchers ran each model three times over 96 total runs, showing output variance can improve vulnerability search by combining different findings from separate passes.
Open models also undercut closed rivals on cost. Three DeepSeek Pro runs cost about $295 and beat Opus 5, Grok 4.6 and Sol, while three DeepSeek Flash runs cost $108 and reached 24 CVEs.
That cheaper coverage carried trade-offs: DeepSeek generated more candidate findings and false leads, while Grok was the most consistent at 21 CVEs found in all three runs despite lower overall recall of 26.
The results suggest open-weight models can now match or surpass publicly available frontier systems in cyber benchmarking, though the report says they still need stronger harnessing and triage to replace premium closed models.