SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
breakthroughsWTF 5.0via arXiv cs.AI

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models

"teaching models to stop spamming useless tool calls like a junior dev on stack overflow"

Explain Like I'm Normal

Researchers developed a new training method for vision-language models that rewards them for finding 'necessary evidence' rather than just guessing the right final answer. By penalizing redundant or off-target tool calls, the model learns to strategically use image cropping and web search to solve complex visual tasks. This moves agents away from 'lucky guesses' and toward actual reasoning-based information gathering.

Read original ↗
#vlm#agentic-ai#tool-use#reinforcement-learning

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.