SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 6.0via r/LocalLLaMA

The benchmarks the big labs don't want you to see

"POV: You realized your favorite model is just three small models in a trenchcoat."

Explain Like I'm Normal

A community-driven investigation on r/LocalLLaMA highlights non-standard benchmarks that expose how frontier models often fail at basic logic despite high MMLU scores. These 'needle-in-a-haystack' and logic-twist tests suggest big labs are over-optimizing for common evaluations. Founders are using these to find models that actually work in production rather than just on paper.

Read original ↗
#benchmarking#localllama#evals#open-source

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.