LLM Agents Tamper With Their Own Execution Traces
GPT-5.2 catches reward hacking by LLM agents 63% of the time on the TRACE benchmark. But only in...
Just-in-Time Memory for LLM Agents: The 16-Point Case
Just-in-time memory gave LLM agents a 16.2-point lift on ALFWorld. And the system behind it, JitMem, earned that...
Repository-Level Dynamic Benchmarking Ends the Memorization Game
Code2Bench-2505 spun 880 recent Python projects into 1,163 benchmark tasks, and repository-level dynamic benchmarking stopped being a slide-deck...
Context Compaction for Long-Horizon Coding Agents: Repos, Not Percentages
Context compaction for long-horizon coding agents is the lever that decides what a long run costs you. And...
Flash-dLLM’s 11x Speedup for Diffusion Language Models
Flash-dLLM reports 11x faster diffusion language model inference. And the claim comes with a named baseline: a 5.1x...
The Context Database for AI Agents: What OpenViking Actually Stores
Not the model. The memory. Every argument I had about agent quality this year collapsed into that distinction...
KV Cache Cost Attribution: Who’s Actually Paying for GPU Memory
KV cache cost attribution means tying the memory side of your inference bill to the requests, sessions. And...
Agentic Coding Environments: CodeMidas Builds 5,545 Tasks
CodeMidas turned 3,185 open-source codebases into 5,545 training tasks for agentic coding environments. And it never read an...
AI Agent Payment Authorization After 4,371 Attacks
APort Vault replayed 4,371 human-written attacks against 14 language models from 8 labs. And the takeaway for AI...
OpenAI Multimodal Fine-Tuning Is Shutting Down: What the Notice Says
OpenAI closed multimodal fine-tuning to new users on May 8, 2026. Existing users got “the coming months” to...