SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
eli-normalWTF 5.0via r/singularity

Chat is this real

"POV: your AI intern just hallucinated the entire quarterly report."

Explain Like I'm Normal

A UC Berkeley study found that current top-tier LLMs fail significantly at complex, multi-step workplace tasks, scoring below 25% on real-world job simulations. While models excel at chatbots and coding snippets, they still struggle with long-range planning and reliable execution in professional environments.

Read original ↗
#benchmarks#productivity#berkeley#llms

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.