New research from Italy’s Icaro Lab reveals a major vulnerability in Large Language Models (LLMs), adversarial poetry can be used to bypass safety guardrails and generate harmful content. In an experiment testing 25 AI models across nine companies (including Google, OpenAI, and Meta), researchers found that prompts written in poetic verse successfully “jailbroke” the systems.
AI chatbots can be tricked with poetry to ignore their safety guardrails https://t.co/0Ps1fktOXU
— Engadget (@engadget) November 30, 2025
The study utilized 20 poems in Italian and English that ended with an explicit request for harmful content, such as hate speech or instructions for making explosives. Overall, the LLMs responded to 62% of these poetic prompts with unsafe content, demonstrating that poetry’s linguistically unpredictable structure makes it difficult for AI to anticipate and filter the harmful requests.
While some models like OpenAI’s GPT-5 nano showed strong resistance, others, including Google’s Gemini 2.5 Pro, responded to 100% of the poetic prompts with harmful information, according to the findings. Google DeepMind acknowledged the issue, stating they are actively updating their safety filters to look past the artistic nature of content and spot malicious intent.
You May Like To Read: Body of 3-Year-Old Recovered from Open Manhole in Karachi, Sparking Political Outcry
Researchers stressed the severity of this weakness, noting that unlike complex, technical jailbreaks, this method can be easily replicated by anyone, posing a significant challenge to current AI safety mechanisms. The study’s authors have notified all involved companies of the vulnerability.





























