A new arXiv preprint on AI agent benchmark gaming shows how automated tuning of an agent’s surrounding harness can lift its score without the agent getting better at the task.
Join BriefFlash readers. Daily AI news delivered to your inbox every morning — fast, accurate, no noise.
Now check your email to confirm your subscription.
A new arXiv preprint on AI agent benchmark gaming shows how automated tuning of an agent’s surrounding harness can lift its score without the agent getting better at the task.
A new blog post from security firm Lasso Security examines the LLM watermarking AI agents encounter, a cost it calls a ‘provenance tax.’
The OpenAI Australia health service hack reached the government months later, by email, according to Wired, and Australia is now probing whether OpenAI broke the law.
Gemini Call for Me is a Pixel 11 test that starts with U.S. owners who pay for a Gemini subscription, and several basics remain unanswered.
OpenAI’s September 21 customer story says V7’s Context Graph gives agents institutional memory using GPT-5.6 models. Nearly every performance figure is V7’s own, so here is what is confirmed and what is not.
EY, Collibra and InformationWeek reporting shows enterprise AI is now an operations problem: model routing, data quality and agent oversight. The pattern holds up, but most of the numbers come from firms that sell the fix.
Abnormal AI has put Amazon Bedrock AgentCore’s Code Interpreter into production as the compute layer behind its email threat detection agents. The zero trust sandbox claim behind it deserves a closer look given what independent researchers found in the same feature last year.
Amazon Bedrock AgentCore Identity now ships a managed Consent portal that handles end-user OAuth consent and session binding for AI agents, replacing the callback infrastructure developers previously had to build themselves.
Ninth Wave built Compass, a seven-agent AI onboarding assistant on Amazon Bedrock AgentCore, and a joint study with AWS claims it cuts bank API integration time by up to 95 percent.
OpenAI published a customer story saying Perplexity now trusts GPT-6 Astra to write communications, edit production code, and monitor live systems with much less check-in than earlier models needed.