More than half of AI-generated patches are broken

More than half of AI-generated patches are broken
New research from 1Password found that ChatGPT 5.5 and Claude Opus 4.8 often fail to fully patch high-impact vulnerabilities and can introduce new bugs instead. The findings suggest AI-assisted patching still needs human review, even as models like GPT 5.5, GPT-5.6-Sol, and Anthropic’s Mythos are promoted for stronger security capabilities. #ChatGPT #ClaudeOpus #1Password #Veracode #ProjectGlasswing #Daybreak

Keypoints

  • 1Password tested ChatGPT 5.5 and Claude Opus 4.8 on six high-impact CVEs.
  • The models fully patched the vulnerabilities less than half the time.
  • The “Copy Fail” Linux cloud kernel flaw was among the tested issues.
  • AI patches often fixed only part of the problem or introduced new bugs.
  • Experts say human review is still necessary for reliable code patching.

Read More: https://cyberscoop.com/ai-code-patching-security-risks/