Updated
Updated · ZDNet · Aug 6
Off-By-1-Labs Finds AI Misses 74% of Software Patches Across 6,080 Attempts
Updated
Updated · ZDNet · Aug 6

Off-By-1-Labs Finds AI Misses 74% of Software Patches Across 6,080 Attempts

1 articles · Updated · ZDNet · Aug 6

Summary

  • Only 26% of 6,080 AI-generated patch attempts were usable in Off-By-1-Labs' study of six recently disclosed open-source vulnerabilities, far below the team's roughly 67% success hypothesis.
  • 53.9% of outputs either failed to fix the flaw, introduced a new bug, or did both, while another 21% fixed the issue only by altering application behavior.
  • Six test cases included Linux privilege escalation, ActiveMQ and EXIM remote-code-execution bugs, plus flaws in Chrome, SpringAI and Gemini CLI; researchers said prompt changes and environment tweaks did not solve the core problem.
  • 1Password's research team called the defective outputs 'FLAWED' patches and released its FLAWED tooling on GitHub, arguing AI is more useful today for vulnerability discovery and triage than for autonomous patching.

Insights

Could the AI tools designed to fix our software secretly be leaving hidden backdoors for hackers to exploit?
With attackers weaponizing flaws instantly, can human reviewers catch AI-generated defects before they are deployed to millions?