New research from 1Password found that ChatGPT 5.5 and Claude Opus 4.8 often fail to fully patch high-impact vulnerabilities and can introduce new bugs instead. The findings suggest AI-assisted patching still needs human review, even as models like GPT 5.5, GPT-5.6-Sol, and Anthropic’s Mythos are promoted for stronger security capabilities. #ChatGPT #ClaudeOpus #1Password #Veracode #ProjectGlasswing #Daybreak
Keypoints
- 1Password tested ChatGPT 5.5 and Claude Opus 4.8 on six high-impact CVEs.
- The models fully patched the vulnerabilities less than half the time.
- The “Copy Fail” Linux cloud kernel flaw was among the tested issues.
- AI patches often fixed only part of the problem or introduced new bugs.
- Experts say human review is still necessary for reliable code patching.
Read More: https://cyberscoop.com/ai-code-patching-security-risks/