breakthroughsWTF 6.4via arXiv cs.AI
Rater State Bias in RLHF Preference Data: An Audit Framework
"Your AI is depressed because its underpaid human trainer is having a bad day."
Explain Like I'm Normal
Researchers found that the emotional state and stress levels of human raters significantly contaminate the RLHF data used to train models. This 'state shift' means AI models aren't just learning what a good answer looks like, but are actually absorbing the burnout and distress of their annotators. This creates a hidden, structured bias that traditional denoising methods currently fail to catch.
#rlhf#alignment#bias#preference-tuning
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.