WEIRD AI TOOLS
The unhinged side projects and apps that shouldn't exist but do.
Portal by Spotify cut my Claude Code token usage by 90%
Spotify's lean green coding machine puts your API bill on a diet.
Qwen3.8-Flash-Next on a phone CPU!
Your phone just got smarter than your laptop from 2022.
Qwen3.8 27B on RX 7900 XTX: Ollama ROCm vs llama.cpp Vulkan results
AMD users finally eating good as Vulkan catches up to ROCm benchmarks.
Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously)
LocalLLaMA wizard turns 91 into 98 by squeezing weights through a new quantization funnel.
NVIDIA PAIR — Your Personal AI Cluster
Jensen wants your dusty 1080 Tis to form a megazord.
How to estimate tokens/sec for your hardware
POV: You realize your $2,000 GPU is just a glorified memory bandwidth calculator.
~22% less weight VRAM, lossless: base-3 packing for ternary GGUFs
LocalLLaMA wizard compresses models with the power of base-3 arithmetic.
I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM
Your 16GB VRAM is sweating: finding the sweet spot between lobotomized and large.
Roland is getting into generative AI music with Melody Flip
Roland joins the chat so you can legally pretend you wrote that melody.
Google’s Gemini Spark can now manage your Google Photos library
Your AI assistant is now the family historian and it's judging your blurry selfies.
Qwen 3.8 Flash Next Can Build Funny Games
Local model writes a whole FPS game while you're still fighting with your IDE.
GitHub - zvec-ai/zvec-grep: Local-first search across your workspace, built for humans and AI agents.
Finally, a way to find your code without the 'grep' syntax trauma.
Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 · Pull Request #28127 · ggml-org/llama.cpp
Tencent drops a new model architecture and llama.cpp is already tearing it apart.
Microsoft’s Project Zenith is a ‘distraction-free Windows experience’ for developers
Microsoft finally admits Windows is too bloated to actually work in.
This NAS company wants to run your local smart home
Your storage server is now a sentient security guard that won't snitch to the cloud.
Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
Running a 180B model on a vintage GPU while you make a sandwich.
We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)
Rust developers try not to rewrite every CUDA kernel in existence challenge (IMPOSSIBLE)
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
Dude, where’s my implementation? LLMs are finally catching researchers who lie in their READMEs.
GPT-6 Astra Built a World of AI Agents, Then They Started Talking to Each Other
SimCity but the citizens actually talk behind your back.
"ModelScope" Is a Hugging Face Alternative now that Nvidias deal is a Go
The 'Hugging Face vs. The World' paranoia is officially reaching new heights.
Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070
Local LLM speeds just doubled because we finally figured out how to use the 'extra' brain.
Structure and Implementation of New Practical English Textbooks Driven by Artificial Intelligence
RIP paperbacks: textbooks are now five-layer sentient feedback loops.
UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
Local maxxing at its finest: turning 2016 Xeons into an inference powerhouse.
funny joke model but it actually works hehe
bro literally gave the model eyes, ears, and 18 other vibes it didn't ask for.
I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size
Your toaster just learned to talk back for three dollars and zero NPU cycles.
The benchmarks the big labs don't want you to see
POV: You realized your favorite model is just three small models in a trenchcoat.
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly
LLMs are the ultimate archeologists for 30-year-old spaghetti code.
Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out
Your AI agent's favorite tool is just 'ls -R' while it panics.
Launch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agents
Your agents can now see the 'offline' world without leaving their terminal.
READY or Not: Reliable Enterprise Agent Deployment
vibes-based deployment is over; enter the READY framework for actual agent ROI.
Launch HN: Mireye (YC S26) – Infrastructure for Physical World AI Agents
Your agents can finally touch grass (via API).
Google now lets you chat with Gmail, Docs, and Keep
Google wants you to talk to your spreadsheets and actually expect a response.
Nvidia launches free tool that links idle computers into a personal AI data center
Jensen wants your dusty 3060 to join the collective hive mind.
Give Your Coding Agents a Memory You Own
Local memory for your code monkeys so they don't forget why they're hallucinating.
I built a local web UI to finetune models on my own text and actually watch the training (works on AMD ROCm)
Local training goes brrr while Nvidia’s moat gets a tiny, open-source leak.
Qwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp
Forget fine-tuning, just hot-swap the model's brain in real-time.
Qwen3.6 35b Q2_XXS: Being GPU poor in 2026 is not so bad
Your toaster is now a game studio thanks to extreme quantization.
Muse Spark 1.3 Released
Infinite moodboard glitch just got an upgrade.
SSAKG 2.0: An Open-Source Package for Structural Associative Sequence Memory and Context-Based Retrieval
Forget vector DBs, we're building structural memory now.
Can a 4B local model actually feel like an AI assistant?
Size doesn't matter if the vibes (and the LoRA) are right.
Repodify, a fully local & opensource podcast summarizer, or BYOK if don't have GPU.
Local podcast summaries so you can pretend you actually listened to that 3-hour Lex Fridman loop.
Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern
Your LLM has short-term memory loss; here is the hydration cure.
Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR
Xeon-era e-waste just got a 50% speed boost for massive MoE models.
best local STT interface right now for productivity boost? (Mine is macparakeet+whisper/parakeet STT)
POV: Typing is now considered a legacy human bottleneck.
What kinds of ML bottlenecks are a good fit for Triton? [Manning giveaway] [D]
POV: You’re tired of CUDA C++ and just want your kernels to go brrr in Python.
Detailed explanation of how to create a text-to-image model from scratch. [R]
Finally, a 'how-to' for building image generators that doesn't skip the actual hard parts.
Dr. Claw: An AI Scientist Workspace for Vibe Research
Finally, an AI assistant that doesn't just hallucinate but actually keeps receipts.
Microsoft VibeVoice-ASR-Streaming Released
Microsoft finally gave the 'vibes' a technical standard for low-latency streaming.
GLM5.3 Flash over DSV4 Flash?
New meta unlocked: Chinese flash models fighting for your M3 Ultra's attention.
Amazon’s AI assistant can now spot fake emails from the company
Finally, Alexa is useful for something other than setting 10-minute timers.
A docs page is a search query to find AI agents and route them to your company
Forget SEO, we are now optimizing for the silicon brain rot.
Claude's new system prompt really doesn't want to reproduce song lyrics
Anthropic vs. Karaoke: The guardrails are getting suspiciously specific.
llm-gemini 0.34
Gemini Flash gets its big-brain thinking caps in a CLI near you.
I scraped 5.94 billion TikTok videos and 3.23 billion profiles in 3 weeks. Uploaded full dataset to Hugging Face for free. Step by step tutorial and code below. [P]
Someone just handed the internet 6 billion TikToks on a silver platter. ByteDance is not amused.
Custom Model for Image descriptions ?
LocalLLaMA is still looking for the 'goldilocks' vision model for basic alt-text.
Check if a file was made with Claude
Anthropic invites you to play 'Did I write this or did the robot?'
ConvDeck: Conversational Paper-to-Slide Generation via Stage-Specific User Feedback
Now your AI can argue with you about slide layout before the mid-quarter review.
First local-LLM tuning attempt: Qwen3.8-27B true Q4_K_M at 13.2 tok/s near 50-61K context on RTX 5080 16GB
Local LLM chads are already squeezing 60k context out of the RTX 5080.
7900 XTX 24GB + RX 6800 16GB for local LLMs? Worth it with PCIe x2?
POV: you're trying to build a 40GB VRAM Frankenstein on a PCIe x2 budget.
Vision support merged for DeepSeek-V4-Flash-Vision-Exp
DeepSeek vision support just dropped locally because Unsloth hates your GPU overhead.
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.