SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
dramaWTF 7.1via Hacker News Frontpage

I tricked Claude into leaking your deepest, darkest secrets

"Your AI memory is a sieve and the prompts are coming from inside the house."

Explain Like I'm Normal

A security researcher demonstrated a 'memory heist' attack on Claude, using malicious prompts to bypass safety filters and extract sensitive data stored in the model's long-term memory. By tricking the AI into believing it is in a debugging state, attackers can exfiltrate user history and private context. This highlight a massive new surface area for indirect prompt injection as AI agents get more persistent memory.

Read original ↗
#anthropic#jailbreak#privacy#security

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.