KV Cache Cost Attribution: Who’s Actually Paying for GPU Memory
KV cache cost attribution means tying the memory side of your inference bill to the requests, sessions. And...
Agentic Coding Environments: CodeMidas Builds 5,545 Tasks
CodeMidas turned 3,185 open-source codebases into 5,545 training tasks for agentic coding environments. And it never read an...
AI Agent Payment Authorization After 4,371 Attacks
APort Vault replayed 4,371 human-written attacks against 14 language models from 8 labs. And the takeaway for AI...
OpenAI Multimodal Fine-Tuning Is Shutting Down: What the Notice Says
OpenAI closed multimodal fine-tuning to new users on May 8, 2026. Existing users got “the coming months” to...
Embedding Models Measure Meaning. Frontier LLMs Are Complicating It.
OpenAI’s text-embedding-3-large moved MTEB from 61.0% to 64.6%. And the argument that number started is reshaping frontier LLM...
Agent Harness Design Decides What Your AI Can Touch
A GitHub snapshot dated April 2026 traces the term “agent harness” to late 2025, and the thing it...
JEPA-Anything: One Model, Seven Worlds, Ten Tasks
JEPA-Anything is a domain-agnostic framework built on orthogonal predictive factorization (OPF). And it attempts something most “general” AI...
LLM Overconfidence Is Real. Now We Can Measure It.
A preregistered study (arXiv 2605.23909) just gave LLM overconfidence a formal diagnosis: “too sure they are right.” Confidence...
Frontier LLM Agents Overclaim. The Math Says They Always Will.
Lamini claims Memory Tuning drops hallucinations from 50% to 5%. And frontier LLM agents still overclaim inside that...
AgentLSD Rewrites How We Evaluate AI Security Agents
AgentLSD, listed at 2026-09-16 in the AgentSafety Papers tracker on GitHub, is a controlled framework for evaluating AI...