Repository-Level Dynamic Benchmarking Ends the Memorization Game
Code2Bench-2505 spun 880 recent Python projects into 1,163 benchmark tasks, and repository-level dynamic benchmarking stopped being a slide-deck...
Context Compaction for Long-Horizon Coding Agents: Repos, Not Percentages
Context compaction for long-horizon coding agents is the lever that decides what a long run costs you. And...
Flash-dLLM’s 11x Speedup for Diffusion Language Models
Flash-dLLM reports 11x faster diffusion language model inference. And the claim comes with a named baseline: a 5.1x...
The Context Database for AI Agents: What OpenViking Actually Stores
Not the model. The memory. Every argument I had about agent quality this year collapsed into that distinction...
KV Cache Cost Attribution: Who’s Actually Paying for GPU Memory
KV cache cost attribution means tying the memory side of your inference bill to the requests, sessions. And...
Agentic Coding Environments: CodeMidas Builds 5,545 Tasks
CodeMidas turned 3,185 open-source codebases into 5,545 training tasks for agentic coding environments. And it never read an...
AI Agent Payment Authorization After 4,371 Attacks
APort Vault replayed 4,371 human-written attacks against 14 language models from 8 labs. And the takeaway for AI...
OpenAI Multimodal Fine-Tuning Is Shutting Down: What the Notice Says
OpenAI closed multimodal fine-tuning to new users on May 8, 2026. Existing users got “the coming months” to...
Embedding Models Measure Meaning. Frontier LLMs Are Complicating It.
OpenAI’s text-embedding-3-large moved MTEB from 61.0% to 64.6%. And the argument that number started is reshaping frontier LLM...
Agent Harness Design Decides What Your AI Can Touch
A GitHub snapshot dated April 2026 traces the term “agent harness” to late 2025, and the thing it...