Context Language Models Rewrite Their Own Memory
Context language models are language models that manage their own context natively. They treat working memory as a...
Hacktoberfest 2026 Backs Open-Source AI With 300+ Events
Deadline first, motivation second. That is the usual order for anyone who keeps meaning to get hands-on with...
Multimodal Flow Modeling Just Cleared 82.8 on 150B Tokens
MF-1 posted 82.8, and multimodal flow modeling just moved out of my “huh, interesting” folder and into my...
The New Long-Context Benchmark: Frontier AI Agents Stall At 68%
The new benchmark capped frontier AI agents at 68% accuracy. And that number should change how you pick...
LongHarness Benchmark: 68% Is The Ceiling
LongHarness Bench landed on arXiv September 29, 2026, and the best score in the whole paper is 68%....
KV-streams Cut Agentic RL Training Up to 5x
KV-streams landed on arXiv September 29, 2026, with a blunt claim: training agents that run long stops choking...
Meta Muse on AI Glasses: The Launch Sellers Keep Misreading
Phone version of Meta Muse already live. Glasses version queued for “the coming months.” Between those two states...
Self-Supervised Confidence Training Teaches Reasoning Models When to Stop
Qwen2.5-Math-7B posted a reported +20.10% accuracy gain on AIME2024, and self-supervised confidence training is what earned it. The...
AgentWorld Grades Multi-Agent LLM Teamwork in a 2D RPG
AgentWorld grades multi-agent LLM teamwork inside a persistent 2D RPG. And that framing fixes a real gap in...
Dream-RSI Recursive Self-Improvement Improves the Search, Not the Model
Google DeepMind’s Dream-RSI claims 162x fewer discovery-agent calls, and the trick is not a smarter model. It’s cheaper...