SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
breakthroughsWTF 6.9via arXiv cs.AI

ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models

"LLMs can't say no to a spicy bomb recipe if it's presented as an 'artistic critique.'"

Explain Like I'm Normal

Researchers discovered a new 'ASCII Attack' that bypasses safety filters by wrapping harmful instructions inside ASCII art characters. Because the model interprets the prompt as an aesthetic request for feedback rather than a direct command, it bypasses standard alignment training that focuses on plain text surfaces.

Read original ↗
#red-teaming#llm-security#jailbreak#ascii-art

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.