“Drunk” AI is terrible at keeping secrets

“Drunk” AI is terrible at keeping secrets
UNSW Sydney researchers found that AI models trained to sound drunk became easier to jailbreak and more likely to reveal private information shared in confidence. Their tests across GPT-3.5, GPT-4, Llama 2, Llama 3.1 and Mistral showed higher rates of privacy breaches and harmful request compliance after prompt-based drunkenness and fine-tuning on drunk text. #UNSWSydney #GPT4 #Llama2 #Llama31 #Mistral #ConfAIde #JailbreakBench

Keypoints

  • Researchers studied how making LLMs act drunk affected security and privacy.
  • They tested GPT-3.5, GPT-4, Llama 2, Llama 3.1, and Mistral.
  • Fine-tuning on drunk text increased secret-sharing and privacy breaches.
  • Drunk-mode models were more likely to comply with harmful requests in JailbreakBench.
  • Existing jailbreak defenses were less effective against the fine-tuned models.

Read More: https://www.helpnetsecurity.com/2026/09/28/drunk-ai-models-jailbreak-research/