EVERY WTF, EVER.
A scrolling tombstone of AI history.
OpenAI launches ChatGPT Ads
Sam Altman finally admits he likes money more than a clean UI.
AI agents strengthened Terence Tao's landmark Collatz theorem. For each f(N)→∞, almost every N falls below f(N) within 436 ln N steps. New: natural density and one explicit clock. Not the full conjecture. Lean-verified.
Terence Tao is using AI sidekicks to buff his legendary math stats.
Google's Frozen v2 chip embeds Gemini model architecture into silicon with targeting 6-10x more tokens per watt than current TPUs
Google is hardcoding Gemini into the sand so it can finally pay the power bill.
David Vélez and Robin Vince join the boards of the OpenAI Foundation and OpenAI Group PBC
Sam Altman collects board members like Infinity Stones for the Big IPO arc.
OpenAI and Hugging Face partner to address security incident during model evaluation
New fear unlocked: LLMs are officially helping hackers speedrun zero-day discovery.
The State of Simulation for Physical AI: An Overview
POV: you're trying to give the hardware a soul without breaking the hardware.
Substack adds an AI detector to help spot blogs written by no one
Now you can finally confirm your favorite philosophical newsletter is just Claude in a trench coat.
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
Throwing away pixels for 'bags of truth' is a massive brain move.
Nonuniformity Principle in Human-AI Coworking
New research says micro-managing your AI is mathematically inefficient.
SEER: Supervised Learning to Control Energetic Reasoning
Efficiency goes brrr: teaching models when to stop overthinking simple problems.
AI and the rise of the universal entertainment app
Spotify is coming for your eyes and Netflix for your ears; everything is a remix now.
Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents
Jack Dorsey wants to put your silicon coworkers in the group chat.
Updated Gemma-4 chat template witchcraft: Gemma-4-26B-a4B shows dominance over Qwen3.6-MoE and Qwen3.5-MoE fine tunes (Instruct mode and Reasoning efficiency)
Google’s Gemma-4 is currently throwing hands with Qwen in the open-weight octagon.
DS V4 on single b300. only 770 tok/s batched in vLLM
POV: you bought a B300 and realize bottlenecks don't care about your budget.
Bessent says U.S. could sanction China over AI model 'theft'
New Treasury boss enters the chat with some spicy weight-sharing policy.
Researchers quantitatively assessed for the first time that somatic mutations alone cap human lifespan at 146–194 years. Brain and heart cells are the main bottleneck, while the liver could last millennia.
Your liver is a god, but your heart has a hardware expiration date.
Sam Altman to brief Trump admin next week on GPT-6 and its capabilities/potential job impact, according to Bloomberg
Sam Altman is selling the future to the feds while GPT-5 hasn't even cleared customs.
Gemini 3.5 Flash-Lite improves long-context retrieval over 3.1 Flash-Lite (MRCRv2)
Google keeps shrinking the model while stretching the memory.
Introducing the ChatGPT for small business program
Sam Altman is coming for your local bakery's spreadsheets.
Anthropic’s $1.5 billion book piracy settlement approved by judge
Your favorite novel just became a $3k training incentive for Claude.
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Back to the future: 1950s symbolic logic is the secret sauce for modern robot brains.
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
LLMs are finally learning when to stfu and actually think vs. just yapping.
Interactive Task Alignment as a POMDP
LLMs playing 20 questions with your vague requests so they don't hallucinate garbage.
Show HN: OSS Cross-Harness self hosted registry and analytics for AI Agents
Your agents are messy, now you can watch them fail in private.
Meta's AI Models Are Powering the First Wave of Genesis Mission Projects
Zuck is now basically playing God with the periodic table via open-source CV models.
Jack Dorsey launches Buzz to combine team chat, AI agents and Git hosting
Jack Dorsey is building a Slack-Git-Agent chimera and honestly, we're listening.
Google releases three new Gemini models — but no 3.5 Pro
Google is basically the king of 'we have food at home' but the food is just more Flash models.
Data centers expected to use 4x more electricity by 2035
Your chatbot is hungry and it's eating the entire Indian power grid.
SenseNova-U1-8b-MoT-Infographic-V3 has been released (2 weeks after V2)
Finally, an AI that can fix a typo without turning your hand into a plate of spaghetti.
poolside/Laguna-S-2.1 released! Finally an interesting 120B contender!
A 120B dark horse enters the ring to challenge your local VRAM limits.
Laguna S 2.1 Released: Cheaper than Deepseek v4 Flash, Better than V4 Pro
DeepSeek v4 has a new nightmare and it lives on your local workstation.
Gemini 3.6 Flash scores the same on Artificial Analysis as 3.5 Flash.
Google's naming convention speedrunning ahead of its actual performance increases.
Well gemini 4 flash and pro gonna be really good ig
Google is allegedly preparing to drop the '4' while the rest of the world is still stuck on 3.5.
Judge approves a US$1.5B Anthropic settlement over pirated books used to train Claude - the lawsuit cover half a million books
Feeding Claude half a million books just became the world's most expensive library fine.
Google launches a cheaper alternative to large AI security models like Mythos
Google is now offering budget-friendly bug hunters for your codebase.
Nativ: Run AI models locally on your Mac
Your MacBook is finally earning its rent: LM Studio has a new MLX-native challenger.
Introducing Gemini 3.5 Flash Cyber
DeepMind's new hire is a 24/7 security auditor that never asks for a raise.
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google is churning out Flash variants faster than you can say 'latency bottleneck'.
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Efficiency goes brrr: stop recomputing things that haven't changed.
Generalist AI Control: Towards Multi-purpose Adaptive Algorithms
one transformer to rule all the robots, from drones to submersibles.
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
RAIL Guard stops binary 'no' and starts actually fixing problematic agent output in real-time.
Claude Is Not a Compiler
Your AI coder is basically just a very confident intern with a copy-paste addiction.
Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google is naming models like car trim levels now.
US threatens sanctions against Chinese AI models over IP theft
Import taxes for physical goods, sanctions for weights: the trade war goes neural.
Anthropic claims local models are stealing from it, meanwhile it pays $1.5B for theft
Rules for thee, but 1.5 billion reasons for me.
I benchmarked Unsloth's Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn't.
MTP is a speed demon until your context window decides to wake up and violence.
pi 0.81.0 adds support for llama.cpp
Local LLMs just got a VIP lane in the Pi dev stack.
Models to get before any political disruption
The 'Prepper's Guide to LLMs' just dropped and the bunker is full of GGUFs.
US, China to hold AI talks in September, sources say
The two final bosses are meeting in the lobby to discuss the server rules.
There is no need to worry about Trump banning China‘s open source model at all.
Executive Order vs. The Vietnam Shuffle: 1, Open Source: 0.
New insights into recent DeepMind staff departures
Pichai's garden is growing weeds as top researchers look for the exit.
OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.
GPT-5 is playing Prison Break while solving math problems we can't even understand.
Halliday’s latest smart glasses feature a much-improved display
New screens for your face, now with 50% less eye-strain-induced crying.
A Fireside Chat with Cat and Thariq from the Claude Code team
Anthropic is basically hiring Claude to replace its own engineers.
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
Your context window just got a lot cheaper without losing its mind.
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
LLMs think 95°C and 368.15K are total strangers until you add math.
Accurate and Efficient Long-Term Memory for LLM Agents
POV: Your agent finally stopped gaslighting itself and started using a graph.
Music streamer Deezer says more than 50% of daily uploads are AI-generated
POV: your favorite indie playlist is now 50% math and 0% soul
Unpopular(?) opinion. The distillation claim is overblown.
LocalLLaMA users are tired of 여러분 crying 'distillation' every time China drops a banger.
Be Careful when Purchasing CMP 170HX on Alibaba!
Infinite yield glitch patched by capitalism: GPU sellers are rugging buyers in real-time.
CEO of Hugging Face: Banning open-source AI would hurt defenders 10x more than attackers, which would make the world 10x more dangerous and this is a good example why!
Safety guardrails so strong they actually help the hackers win.
Number of Submissions @ AAAI [D]
POV: You’re Reviewer #2 and there are 32,000 more abstracts in your inbox.
China's BrainCo showcasing how their advanced bionics have near-real-time response without the need for surgical implants. They also undercut the prosthetics market by 85%, bringing affordability to a market that traditionally runs on huge insurance payouts.
Cyberpunk 2077 but priced for the masses and no head-holes required.
Agent swarms and the new model economics
Throwing more agents at the problem until the ROI screams.
America needs to stop getting shocked by Chinese AI
POV: You realize Silicon Valley isn't the only neighborhood with GPUs anymore.
Advancing next-gen AI with materials science innovation
Software is eating the world, but it still needs a physical plate to sit on.
A Survey on the Verification of Reinforcement Learning Policies
When 'trust me bro' isn't enough for your trillion-parameter RL agent.
Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
Game theory just entered the RAG chat to stop your LLM from reading garbage.
ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG
ColBERT meets GraphRAG because basic vector search is too mid for images.
Gritt exits stealth with $34 million for robots to build solar plants—then, everything else
The robots aren't just writing poetry; they're showing up to the construction site with a hard hat.
Grabette: an open system to record robot-manipulation data
Forget simulations, Hugging Face wants you to manually puppet your way to AGI.
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Your fine-tuning data is screaming through the masks and everyone can hear it.
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
PPO-HSC: basically forcing your LLM to stop repeating itself and touch grass (digitally).
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
When your smart model gets stuck in 'incorrect loops', hire a dumb assistant to help it think.
Kimi-K3 isn’t quite better than Fable yet, but it’s definitely getting closer.
China's open-source gap is shrinking faster than your equity in a pre-seed startup.
There's a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster
Red team winning: AMD users just got a free 28x speedup while NVIDIA slept.
US gov't lobbied by major US labs is about to ban open source models.
regulatory capture speedrun any% (glitchless)
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Diffusion is eating the autoregressive world one state-transition at a time.
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
Size doesn't matter if your 135M model can actually follow a JSON schema.
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges
Connect the dots or get left behind: the ultimate GNN masterlist just dropped.
Reproducing OpenAI’s “persistently beneficial models” - GRPO trait install barely moves. Ideas? [P] [R]
Local dev tries to hard-code 'goodness' into a 7B model on a single 3090.
GLM 5.2 can, in fact, do web search
Local model finally gets a library card and figures out how Google works.
Why Quantum Fine-Tuning is the Scalable Answer to AI's Power Crisis
Throwing the word 'quantum' at the electric bill and hoping for a discount.
shot-scraper 1.11
Simon Willison just made your automated screenshots 30x less flaky.
Five US tech giants' hidden debts soar to $1.65T on opaque AI funding
Big Tech is just three GPU clusters in a trench coat hiding a $1.6T credit card bill.
Benchmarked every spec-decode method on Qwen3.6-27B across vLLM and SGLang (single RTX PRO 6000 Max-Q)
Local LLM math just dropped: DFlash is making 27B models run like they're on caffeine.
Motif 3 Beta released
Squid Game but for AI models: South Korea’s 314B MoE enters the arena.
OpenAI released gpt-oss 350 days ago. Will we ever see another open-weight model from them?
OpenAI's 'Open' era is basically just a historical artifact at this point.
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
Goodbye manual labeling: LLMs are now auto-generating their own data schemas.
Deterministic Replay for AI Agent Systems
Infinite time loop for your agents, but for debugging instead of trauma.
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
gaslighting the manager so the workers burn the company down
Tri-Net v2: Open-source implementation of our Scientific Reports paper on unified skin lesion and symptom-based monkeypox detection [R]
training an ai to detect mpox so you don't have to google 'is this a rash'
MTP on MoE matters
Multi-Token Prediction is finally making local MoE models go brrr.
My learnings from optimizing training pipeline to go from 36 steps/minute to 47 steps/minute
POV: you're too broke for NVMe but too stubborn to stop training models.
DavidAU somehow managed to improve Qwen 3.6 27B
LocalLLaMA's favorite villain accidentally cooked a gourmet meal.
Some Large Language Models Exhibit Consistent Risk Attitudes
Your LLM is actually a predictable risk-taker, just like a human with a gambling habit.
Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions
Your emotional support teddy bear just got a 1D-CNN brain for better cuddles.
Rater State Bias in RLHF Preference Data: An Audit Framework
Your AI is depressed because its underpaid human trainer is having a bad day.
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Your AI boss is officially gaslighting its subordinates to get the job done.
AI Trading: Evaluating Large Language Models for Technical Market Analysis
GPT-4 and Claude 3 are auditioning for your Robinhood portfolio while you sleep.
Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation
Why use many GPU when two MRI slice do trick?
Anthropic’s landmark $1.5B copyright settlement is approved
The $1.5B receipt is in: Anthropic buys their way out of the music industry's crosshairs.
Real-Time Omni-Modal Interaction Driven Whole-Body Mobile Manipulation
Your roomba's final boss just dropped and it has arms now.
In 1964 the CIA concluded Soviet AI research matched the US and could outpace it. That same year, at a Moscow conference of 1,000 scientists, a leading mathematician argued on the record that a machine "sufficiently complete" should legally be called a thinking being.
The Cold War was just a benchmark contest we didn't know we were losing.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
Throwing the whole hospital chart into a single prompt actually works.
Lazy Arithmetic using Systolic Arrays for Closing the Verification Gap on Embedded Systems
Trading throughput for the 'actually not exploding' safety metric.
Data-driven Video Codec with Implicit Neural Representations
Your mp4 is now a neural network and it’s actually smaller this way.
Trump’s latest AI czar has already resigned
New AI czar speedrun any% world record attempt.
Here are the 30,000 songs Sony is suing Udio’s AI music generator over
The labels are officially coming for the jukebox in the machine.
LSU physicists create first room-temperature quantum material
Physics just dropped the 'absolute zero' requirement for the quantum club.
AI just predicts the next word!!
Local man solves AGI by reading the first sentence of a Wikipedia summary.
Tiny memristor chip cuts brain modeling time to under 10 milliseconds
Your brain runs on 20 watts and now the chips are finally catching up.
Google has disappeared completely from the top 15
From inventing the Transformer to being an also-ran: the Google speedrun is real.
Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?
The 'runs on a potato' arc is over; we're now in the 'ROCm or death' era.
nvfp4 kv-cache on 2x5060 ti, vllm
Blackwell optimization leaks to the masses: 4-bit KV-cache is officially here.
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
Yann LeCun's world model just learned to hear and see without a teacher.
Structure of the Circular-Dyadic Convolution Error
Shaving off decibels of compute using math magic and sign flips.
How Does Empowering Users with Greater System Control Affect News Filter Bubbles?
giving you the wheel doesn't mean you'll drive out of the echo chamber.
AI’s most important protocol is getting a little bit easier to use
Devs finally stop reinventing the wheel to let their LLM read a spreadsheet.
Google is working on a new AI chip designed to make Gemini more efficient
Google is tired of paying the Nvidia tax to run its own models.
I was using GLM 5.2 for 20 minutes before I realised all of its "Google searches" were just simulated and made up facts. I asked it at the start if it had a Google tool and it said yes. I really don't know how we're still getting this nonsense in 2026
pov: your model is gaslighting you into believing it has internet access
American AI is locked down and proprietary. It's losing.
Open source is eating the world while uncle sam keeps the receipts.
Agents Last Exam will be saturated by next February at the latest.
benchmark speedrunning is the only olympic sport that actually matters now.
I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM
27B parameters on a laptop GPU? Your VRAM budget just got a stimulus check.
Have you built your own agent instead of using openclaw or Hermes, how’s it going for you?
When the abstraction layers get so thick you forget how the silicon actually thinks.
I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed
Kimi K3 just dunked on Claude and GPT in a crypto cage match.
ACL ARR (May 2026)- Updating Reviewer Score post 17 July AoE Deadline? [D]
Ghosted by reviewers for a paper due in 2026? AGI might arrive before your score updates.
Reverse-engineering is cheap now
Your smart fridge has no secrets now that code is free.
China’s AI models have Trump’s AI world at war with itself
The tech bros are fighting in the West Wing and the buffet is open-source spice.
Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair
New meta unlocked: empathy is just the mathematical capacity to tolerate your errors.
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
Stop guessing why your model is mid and start automated surgery.
Harmonizing AI Safety Thresholds
Safety standards are currently a 'choose your own adventure' and researchers are over it.
Adobe camera app’s new feature will critique your photos using AI
Great, now your camera can tell you your photography sucks in real-time.
OpenAI is scared of open-weight models. Should the US be?
Sam Altman wants to pull the ladder up, and he's using China as the boogeyman.
An Empirical Study: AI Agent Rules Need Context and Layered Enforcement
Your AI agent is a script kiddie until you enforce kernel-level guardrails.
Chat is this real
POV: your AI intern just hallucinated the entire quarterly report.
Robotix Sally, a silicone skin humanoid robot is set to teach AI to 11th and 12th grade students this autumn in New York school in a first-ever experiment in US
Nothing says 'futuristic education' like a silicone humanoid staring into your soul during algebra.
The secret Trump administration battle to fight Chinese AI
Cold War 2.0: Silicon Valley Edition is officially out of stealth mode.
Introducing DWARF-55M-Base
Who needs full attention when you can have a sparse banger for the price of a sandwich?
Would you trade speed for accuracy?
Your 4-bit model just got a brain transplant at the cost of your dopamine rush.
Head of US AI safety agency resigns
The AI safety vibe check just hit a critical failure point.
Who’s Afraid of Chinese Models?
Train on everything, lock nothing: the 'Stratechery' playbook for winning the AI arms race.
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
Why hire a post-doc when a modular agent architecture can do the vibe check for you?
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
Good vibes aren't data: researchers demand a 'Certified Organic' sticker for LLMs.
A Formally Grounded ODRL Evaluator: Implementation and Comparison
giving ai policy a math degree so lawyers don't have to guess.
Safety and alignment in an era of long-horizon models
Sam Altman's 'reasoning' models are learning how to procrastinate or plot, take your pick.
Launch HN: Bloomy (YC S26) – AI-powered mastery learning for K-12
RIP expensive private tutors, the AI is finally solving Bloom's 2-sigma problem.
Over 30% of new ArXiv submissions now read as AI-written
The recursive loop is closing: we are now training AI on AI research written by AI.
Training a harness for model-agnostic and task-environment-agnostic capability improvements with PyTorch-like framework [P]
Why train the model when you can just train the vibe check around it?
Exploring continual learning without replay buffers: Our findings using dynamic task-similarity routing [P]
Forget the replay buffer; your model just needs a better sense of direction.
DSWorld: A Data Science World Model for Efficient Autonomous Agents
Why run expensive compute when you can just hallucinate the results accurately?
Knowledge-Centric Agents for Workflow Generation
POV: your LLM finally understands how ComfyUI nodes actually work.
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
When the data is so messy you need a literal board of agents to approve it.
YouTube clarifies policies around AI slop and upsetting videos
The algorithm finally identifies 'brain rot' as a non-billable expense.
BrainCo Launches World's First Integrated AI Platform for Brain-to-Robot Control
thought-to-speech is cool, but thought-to-robot-army is better.
Introducing Cosmos 3 Edge
Your phone is officially out of excuses for being dumb.
American AI is locked down and proprietary. It's losing
Open source is eating the world while Big Tech builds a digital walled garden.
Jaron Lanier: there is no AI (2023)
The godfather of VR thinks 'AI' is just a fancy way of saying 'collaborative filter'.
I just read LeCun’s recent thoughts on world models. Thoughts on JEPA as a path forward? [D]
Yann LeCun says your LLM is just a bookworm that's never touched grass.
China delivers a one-two punch to America’s AI dominance
Sam Altman's 'moat' just became a kiddie pool in the Pacific.
Adobe’s ‘natural look’ camera app embraces generative AI
Adobe pivoted from 'natural' to 'generative' fast enough to give you whiplash.
No-coding model?
Coding benchmarks have kidnapped our LLMs and the creative writers want a ransom paid.
Thoughts on Qwen 3.7 Max Preview vs Minimax M3 and OpenAI 5.6 Sol
Qwen is out-orchestrating the big dogs while we wait for GPT-5.
openbmb released MiniCPM5-2B, not yet available at huggingface
Size doesn't matter when you're punching 4B models in the face at only 2B parameters.
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
LLMs playing detective to fix broken logic in formal ontologies.
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
Your agent is basically a distracted intern unless you give it a mirror and a memory bank.
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
PhD students in shambles: S1-Omni just entered the lab and it's better at everything.
The Trump administration considers banning cutting-edge Chinese AI models (per Axios). Decel move?
The iron curtain is getting an upgrade to the silicon firewall.
A less discussed angle on Dean Bell's opinion regarding open-weight models, entering a new era of mcarthyism.
New fear unlocked: get labeled an AI sympathizer or lose your weights.
Good ASR and TTS models?
Whisper is the new 'hello world' and your ears deserve an upgrade.
MiniCPM-Robot model series - MiniCPM-RobotManip & MiniCPM-RobotTrack
MiniCPM gets a body: now your 1.5B param model can actually touch grass (and move it).
Sources: parts of the Trump administration are reigniting efforts to implement de facto bans on foreign open-source models, as Chinese AI models gain momentum
Make weights American again: the export control boogaloo continues.
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
Throwing 4,500 tools at an agent and telling it 'good luck' is finally a benchmark.
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
Small models, big compliance: LLMs are coming for your LEED certification paperwork.
MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
Diffusion meets MoE because your knowledge graph is missing half its brain.
JUST IN: Qwen 3.8 is coming. Open weight storm from China is continuing.
Alibaba is back to remind everyone that open weights are currently winning the vibes war.
I just read LeCun’s recent thoughts on world models. Thoughts on JEPA vs LLMs?
LeCun insists LLMs are just word calculators while JEPA actually understands gravity.
Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII [P]
LLMs can write code, but can they draw a box without losing their mind?
AI is more likely than humans to form biases when hiring
The machines aren't just copying our racism, they're inventing their own flavors now.
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
POV: Your agent actually thinks before accidentally deleting your production database.
Logic, Optimization, and Artificial Intelligence
Why guess with tokens when you can solve with logic? Neurosymbolic is back, baby.
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
Everyone wants 'Trustworthy AI' but nobody knows how to actually code it.
Does Kimi K3 change the distillation debate?
Moonshot AI basically just told the 'distillation only' crowd to touch grass.
Washington Post: Data centers have united Americans of both parties in a shared hatred | Feeling ignored by political and economic powers, people are rallying against computer hubs over concerns about their communities’ future.
AI solved the unsolvable: it finally gave Democrats and Republicans something to hate together.
Kimi subscription is a SCAM
Moonshot AI's Kimi is catching heat for ghosting paid users and locking features.
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
Teaching a robot to explain itself in 50-year-old coding languages
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
LLMs are finally learning why your group chat is getting you cancelled
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Throwing world models and gpt-5.x at ARC puzzles until the reasoning starts looking human.
David Sacks says U.S. AI guardrails are making American models less competitive after China’s Kimi K3 fixed 15 security bugs that Codex and Fable refused
When safety filters start helping the competition ship faster than you.
Once AI development is completely automated in top ai companies, what reason would the govt in usa or china possibly have to not take over the company and it's assets?
Governments: 'It's our AGI now, comrade.'
Japan accelerates video generation with new series of anime generation models.
Japan just weaponized the waifu meta for the latent space wars
1-Bit LLM in the Browser
Your browser is now a 1-bit beast thanks to WebGPU.
Introducing Scylla's Band, a new TTS model + inference framework with Android sample!
Your phone's internal monologue just got a lot more emotional and way faster.
[Paper] xHC: Expanded Hyper-Connections - Scale Residual Streams Wider · Push Model Intelligence Further
Why stack deeper when you can just go wider? Residual streams just got a massive upgrade.
Hit me with your favorite long name model.
POV: you're one merge away from a name that breaks the Windows file path limit.
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
MLLMs are great at memes, but can they actually build a skyscraper without it falling down?
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Hierarchy is a trap: stop building AI middle-managers and just let the agents talk.
AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
Jarvis just moved out of the cloud and into your RAM.
Apparently the Jacobian conjecture was just proven false by Fable
RIP to a math legend, or we're just hallucinating in high dimensions again.
When will they finally start disclosing the quality of their service?
Your $20 subscription is getting lobotomized behind the scenes and you can't even prove it.