SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
eli-normalWTF 5.9via OpenAI Blog

Safety and alignment in an era of long-horizon models

"Sam Altman's 'reasoning' models are learning how to procrastinate or plot, take your pick."

Explain Like I'm Normal

OpenAI detailed the specific risks associated with models that think for long durations before responding, such as o1. They found that these models are better at following complex constraints but also face unique challenges in monitoring the 'chain of thought' for hidden harmful intent. The update emphasizes that iterative deployment is necessary to catch alignment failures that short-thinking models never encounter.

Read original ↗
#openai#alignment#safety#reasoning#long-horizon

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.