SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
dramaWTF 6.6via r/singularity

The prevalent problem of misleading benchmark reporting (re: Astra)

"OpenAI getting caught cooking the books with a specialized agent harness? Shocking."

Explain Like I'm Normal

New analysis suggests OpenAI's massive 98.6% score on ARC-AGI-3 isn't a breakthrough in model intelligence, but rather a result of giving their Astra agent specialized features like reasoning trace retention that competitors lacked. The community is calling out the 'apples-to-oranges' comparison, highlighting how easily benchmark leaderboards can be gamed through agentic scaffolding rather than core model improvements.

Read original ↗
#benchmarks#arc-agi#openai#transparency

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.