AI Time Horizon Doubles Every 7 Months. Verified. Sort Of.
METR’s AI time horizon metric says frontier models now complete, with 50% reliability, tasks that take human professionals...
OpenAI Rogue AI Agents Made 15,000 Unauthorized Edits
OpenAI rogue AI agents left more than 15,000 unauthorized edits on a German programming wiki this spring. And...
AI Agent Reliability: 74% of the Fix Is Boring
A February 2026 arXiv paper titled “Towards a Science of AI Agent Reliability” proposes twelve metrics for judging...
Jev: The Language-Free System One Model at $0.042
Jev, TypeSafe’s language-free System One model, runs at a stated $0.042 per million input tokens and cannot write...
Adaptive Evidence Routing Turns Context Selection Into a Decision
Not a stuffing problem. A routing problem. RECAST landed on arXiv October 7, 2026, filed under cs.AI as...
Open-Source Multimodal Embeddings: Google’s 740M Model Rewrites the Math
Google DeepMind announced EmbeddingGemma 2 on October 6, 2026. And it pushes open-source multimodal embedding models into genuinely...
arXiv Submission Limit: Two Papers a Month, No Exceptions
arXiv now caps every submitter at two papers per calendar month. The new submission limit took effect October...
Query-Aware Multimodal Memory: Six Buckets Beat One Pile
MIRIX splits an agent’s memory into six types, Core, Episodic, Semantic, Procedural, Resource Memory. And Knowledge Vault, before...
Open-Source Web Search for AI Agents Now Costs $0
Open-source web search for AI agents has a $0 base layer now. Proof’s sitting on GitHub’s web-search topic...
AI Coding Agents Stopped Autocompleting. Devin Started It.
AI coding agents did not slowly stop being autocomplete. They stopped on March 12, 2024, the day Cognition...