Open-Source Multi-Agent Frameworks: The Ones I’d Actually Ship
Open-source multi-agent frameworks stopped being conference demos, and nobody threw a party about it. Smolagents sits at 29,177...
AI Code Refactoring Tools 2026: The 91% Trap
Claude Code reported 91% refactor accuracy on a 150K-line codebase. And that single figure is both the best...
Prompt Optimization Frameworks: Put a Number on the Prompt
Prompt optimization frameworks have their proof point and it is not subtle: OPRO beat human-designed prompts by up...
GLM-5.3-Flash: The Ox Alpha Reveal, Specs, Pricing, and Open Weights
GLM-5.3-Flash spent the back half of August 2026 answering to a name that wasn’t its own. And if...
Error-Driven Prompt Optimization: Your Failures Are the Training Data
Four steps. That’s all ETGPO needs. Error collection, error taxonomy creation, error category selection, guidance generation. The whole...
Declarative UI Generation At Small-Model Cost: Amazon’s Play
Amazon has five names on an arXiv paper, identifier 2609.04184. And its title states the whole thesis: “Toward...
Latent Reasoning Moves AI Thinking Off Your Token Bill
A 2025 arXiv paper scaled a latent reasoning model to 3.5 billion parameters and 800 billion training tokens....
New Reasoning Image Generation Benchmarks End The Beauty Contest
R2I-Bench fires 3,068 reasoning-heavy prompts at text-to-image models, and pretty can’t save you anymore. It’s a reasoning image...
Claude Fable 5.1 Cut Cache Reads 75%. Here’s the Math.
Start with the cache-read line, not the eval charts. Claude Fable 5.1 is that rare launch where the...
User Simulation Evaluation Frameworks: 53% Against Real Humans
User simulation matched real human choices 53% of the time. That’s the headline result from a 2026 study,...