AI BREAKTHROUGHS
New capabilities, new scaling laws, new ways the future got weirder.
AI agents strengthened Terence Tao's landmark Collatz theorem. For each f(N)→∞, almost every N falls below f(N) within 436 ln N steps. New: natural density and one explicit clock. Not the full conjecture. Lean-verified.
Terence Tao is using AI sidekicks to buff his legendary math stats.
Google's Frozen v2 chip embeds Gemini model architecture into silicon with targeting 6-10x more tokens per watt than current TPUs
Google is hardcoding Gemini into the sand so it can finally pay the power bill.
The State of Simulation for Physical AI: An Overview
POV: you're trying to give the hardware a soul without breaking the hardware.
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
Throwing away pixels for 'bags of truth' is a massive brain move.
Nonuniformity Principle in Human-AI Coworking
New research says micro-managing your AI is mathematically inefficient.
SEER: Supervised Learning to Control Energetic Reasoning
Efficiency goes brrr: teaching models when to stop overthinking simple problems.
Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)
Google’s Gemma-4 is currently throwing hands with Qwen in the open-weight octagon.
Gemini 3.5 Flash-Lite improves long-context retrieval over 3.1 Flash-Lite (MRCRv2)
Google keeps shrinking the model while stretching the memory.
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Back to the future: 1950s symbolic logic is the secret sauce for modern robot brains.
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
LLMs are finally learning when to stfu and actually think vs. just yapping.
Interactive Task Alignment as a POMDP
LLMs playing 20 questions with your vague requests so they don't hallucinate garbage.
Meta's AI Models Are Powering the First Wave of Genesis Mission Projects
Zuck is now basically playing God with the periodic table via open-source CV models.
poolside/Laguna-S-2.1 released! Finally an interesting 120B contender!
A 120B dark horse enters the ring to challenge your local VRAM limits.
Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro
DeepSeek v4 has a new nightmare and it lives on your local workstation.
Well gemini 4 flash and pro gonna be really good ig
Google is allegedly preparing to drop the '4' while the rest of the world is still stuck on 3.5.
Introducing Gemini 3.5 Flash Cyber
DeepMind's new hire is a 24/7 security auditor that never asks for a raise.
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google is churning out Flash variants faster than you can say 'latency bottleneck'.
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Efficiency goes brrr: stop recomputing things that haven't changed.
Generalist AI Control: Towards Multi-purpose Adaptive Algorithms
one transformer to rule all the robots, from drones to submersibles.
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
RAIL Guard stops binary 'no' and starts actually fixing problematic agent output in real-time.
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google is naming models like car trim levels now.
A Fireside Chat with Cat and Thariq from the Claude Code team
Anthropic is basically hiring Claude to replace its own engineers.
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
Your context window just got a lot cheaper without losing its mind.
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
LLMs think 95°C and 368.15K are total strangers until you add math.
Accurate and Efficient Long-Term Memory for LLM Agents
POV: Your agent finally stopped gaslighting itself and started using a graph.
Agent swarms and the new model economics
Throwing more agents at the problem until the ROI screams.
Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
Game theory just entered the RAG chat to stop your LLM from reading garbage.
ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG
ColBERT meets GraphRAG because basic vector search is too mid for images.
Gritt exits stealth with $34 million for robots to build solar plants—then, everything else
The robots aren't just writing poetry; they're showing up to the construction site with a hard hat.
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Your fine-tuning data is screaming through the masks and everyone can hear it.
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
PPO-HSC: basically forcing your LLM to stop repeating itself and touch grass (digitally).
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
When your smart model gets stuck in 'incorrect loops', hire a dumb assistant to help it think.
Kimi-K3 isn’t quite better than Fable yet, but it’s definitely getting closer.
China's open-source gap is shrinking faster than your equity in a pre-seed startup.
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Diffusion is eating the autoregressive world one state-transition at a time.
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
Size doesn't matter if your 135M model can actually follow a JSON schema.
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges
Connect the dots or get left behind: the ultimate GNN masterlist just dropped.
Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]
Local dev tries to hard-code 'goodness' into a 7B model on a single 3090.
GLM 5.2 can, in fact, do web search
Local model finally gets a library card and figures out how Google works.
Motif 3 Beta released
Squid Game but for AI models: South Korea’s 314B MoE enters the arena.
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
Goodbye manual labeling: LLMs are now auto-generating their own data schemas.
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
gaslighting the manager so the workers burn the company down
Some Large Language Models Exhibit Consistent Risk Attitudes
Your LLM is actually a predictable risk-taker, just like a human with a gambling habit.
Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions
Your emotional support teddy bear just got a 1D-CNN brain for better cuddles.
Rater State Bias in RLHF Preference Data: An Audit Framework
Your AI is depressed because its underpaid human trainer is having a bad day.
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Your AI boss is officially gaslighting its subordinates to get the job done.
AI Trading: Evaluating Large Language Models for Technical Market Analysis
GPT-4 and Claude 3 are auditioning for your Robinhood portfolio while you sleep.
Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation
Why use many GPU when two MRI slice do trick?
Real-Time Omni-Modal Interaction Driven Whole-Body Mobile Manipulation
Your roomba's final boss just dropped and it has arms now.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
Throwing the whole hospital chart into a single prompt actually works.
Lazy Arithmetic using Systolic Arrays for Closing the Verification Gap on Embedded Systems
Trading throughput for the 'actually not exploding' safety metric.
Data-driven Video Codec with Implicit Neural Representations
Your mp4 is now a neural network and it’s actually smaller this way.
LSU physicists create first room-temperature quantum material
Physics just dropped the 'absolute zero' requirement for the quantum club.
Tiny memristor chip cuts brain modeling time to under 10 milliseconds
Your brain runs on 20 watts and now the chips are finally catching up.
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Yann LeCun's world model just learned to hear and see without a teacher.
Structure of the Circular-Dyadic Convolution Error
Shaving off decibels of compute using math magic and sign flips.
Google is working on a new AI chip designed to make Gemini more efficient
Google is tired of paying the Nvidia tax to run its own models.
Agents Last Exam will be saturated by next February at the latest.
benchmark speedrunning is the only olympic sport that actually matters now.
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
27B parameters on a laptop GPU? Your VRAM budget just got a stimulus check.
I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed
Kimi K3 just dunked on Claude and GPT in a crypto cage match.
Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair
New meta unlocked: empathy is just the mathematical capacity to tolerate your errors.
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.