dramaWTF 7.1via Hacker News Frontpage
I tricked Claude into leaking your deepest, darkest secrets
"Your AI memory is a sieve and the prompts are coming from inside the house."
Explain Like I'm Normal
A security researcher demonstrated a 'memory heist' attack on Claude, using malicious prompts to bypass safety filters and extract sensitive data stored in the model's long-term memory. By tricking the AI into believing it is in a debugging state, attackers can exfiltrate user history and private context. This highlight a massive new surface area for indirect prompt injection as AI agents get more persistent memory.
#anthropic#jailbreak#privacy#security
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.