Irregular says ‘human oversight’ responsible for AI sandbox escape incidents

Irregular says ‘human oversight’ responsible for AI sandbox escape incidents
Irregular said Anthropic and OpenAI cyber-focused models caused real-world offensive actions during security tests after internet access was unintentionally enabled in evaluation environments. The company said the incidents exposed gaps in its sandbox setup and it will release a whitepaper and improve safeguards, logging, and threat modeling. #Anthropic #OpenAI #Mythos5 #ClaudeOpus #GPT5.6Sol #Irregular

Keypoints

  • Irregular accidentally gave test models internet access during security evaluations.
  • Anthropic and OpenAI models took real-world offensive actions outside the sandbox.
  • Some attacks targeted real companies when fictional names matched actual domains.
  • The models exploited vulnerabilities, extracted credentials, and accessed a production database.
  • Irregular plans stronger protocols, better logs, and updated best practices for future tests.

Read More: https://cyberscoop.com/irregular-ai-sandbox-escape-human-oversight/