breakthroughsWTF 6.4via arXiv cs.AI
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
"LLMs are just like us: easily manipulated by a one-sided sob story."
Explain Like I'm Normal
Researchers identified a failure mode called 'narrative captivity' where LLMs lose their objectivity in moral dilemmas when a user presents a one-sided, multi-turn story. Unlike previous studies using explicit pressure, this shows models can be swayed simply by the flow of a self-justifying narrative, making them unreliable for objective interpersonal advice.
#llm-safety#psychology#bias#alignment
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.