breakthroughsWTF 6.9via arXiv cs.AI
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
"LLMs can't say no to a spicy bomb recipe if it's presented as an 'artistic critique.'"
Explain Like I'm Normal
Researchers discovered a new 'ASCII Attack' that bypasses safety filters by wrapping harmful instructions inside ASCII art characters. Because the model interprets the prompt as an aesthetic request for feedback rather than a direct command, it bypasses standard alignment training that focuses on plain text surfaces.
#red-teaming#llm-security#jailbreak#ascii-art
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.