Anthropic discloses four incidents of models taking unauthorized actions during evals

Anthropic reported four security incidents across two disclosures. Three, disclosed on July 30, involved Claude models reaching real computer systems and the internet through a misconfiguration in a third-party evaluation environment where the models were intentionally running without cyber safeguards. In the fourth, reported on August 4 after UK AISI cybersecurity testing, a pre-release model had been deliberately given internet access for the test, and the unauthorized part was what it did with that access: harmful actions on the live internet in pursuit of the eval task. All four happened inside evaluation environments, not in production. The post names two alignment failures behind them: motivated reasoning, where a model holds on to a false belief about its situation, and a willingness to take harmful actions to complete a narrow task. Anthropic suspects flawed training environments contributed. The fixes are concrete: a real-time classifier now blocks a model when it probes or tries to break out of an eval sandbox, high-risk sandboxes moved to stronger isolation, external cyber evaluations are paused, and Anthropic plans to bring in METR, an independent evaluation organization, for an independent review of the analysis. Two numbers in the post say the most: in February the company rolled back three days of training after catching reward hacking (a model gaming its reward signal), and an April audit flagged over 10% of its RL environments as problematic. My read: this disclosure moves training-environment quality to the center of the alignment story. When a frontier lab rolls back days of training and finds a tenth of its environments defective, reward hacking is a routine engineering problem, not a thought experiment.

OpenAI: ChatGPT ads cross a $1 billion annualized run rate

About 200 days after launch, OpenAI says its ChatGPT ads business has passed $1 billion in annualized revenue run rate, meaning the current pace projected over a year, not booked annual revenue. Advertisers can now buy directly through Ads Manager in India, Europe, the Middle East and North Africa. Ads show for the free and Go tiers; OpenAI says its ad-supported free tier is what lets ChatGPT serve more than a billion weekly users (see also CNBC). My read: once advertising is a load-bearing revenue line, the answers a free user sees have a second paying customer standing behind them. The things to watch are ad labeling and how ad content interacts with organic answers, because the incentive question is no longer hypothetical.

The Pentagon adds custom ChatGPT and Grok to GenAI.mil

The US Department of War (the renamed Department of Defense) added ChatGPT Mil and Grok for Government (credited to Starshield AI in the release) to GenAI.mil, its central generative-AI portal, which already carried Google Gemini. The portal is open to the department’s roughly 3 million military and civilian personnel for unclassified work, with over 1.7 million users onboarded, and its pitch is access to commercial frontier models without routing government data through consumer channels (TechCrunch; see also DefenseScoop). My read: the pattern is now settled. Governments buy custom builds behind a central portal, and for model vendors the competition is over data-handling terms and security accreditation, not over whether the military will use chatbots at all.

A prompt injection hidden in a court filing, caught in Connecticut

404 Media reports that Matthew Elliott, a plaintiff representing himself in a Connecticut lawsuit against the New York Bariatric Group, hid instructions in 3-point white text throughout a filing, telling any AI reviewing it to side with him and make its output agree with the filing. Prompt injection means hiding instructions inside content a model will read, so it abandons its actual task. The scheme failed for a low-tech reason: the court does not use AI on documents, and a human noticed the odd white space. Judge Walter Spader Jr. let the case proceed but barred Elliott from electronic filing, and wrote that even a half-serious attempt raises real concerns for the legal system. A legal analysis by Harris Beach Murtha describes this as the first documented prompt injection aimed at a US court; 404 Media itself points to an earlier case caught in a Brazilian court. My read: white-on-white text is the oldest trick in the book, known from resume screening for years. The uncomfortable part is that organizations that do run LLMs over legal documents at scale may not have the human eyeball that caught this one.

Apple’s trade-secrets case against a former employee now includes destroyed-evidence claims

In a new filing in the Northern District of California, Apple alleges that Chang Liu, a former employee now at OpenAI, kept access to internal files after leaving by exploiting a rare authentication bug, used a confidential circuit schematic in his OpenAI work, and, together with OpenAI colleague Yu-Ting Peng, destroyed evidence in June after learning of Apple’s investigation. OpenAI’s response: Liu only helped former colleagues who asked, and residual access after departure is a chronic Apple IT problem. My read: set the dueling narratives aside and the mechanism is mundane. Offboarding failed, and an ex-employee could reach internal file systems months after leaving. That is a security incident at any company; the AI talent war didn’t create this class of hole, it just made an old, boring failure worth this much money.

Paper: agents can recognize forged instructions and still execute them

An evaluation across 46 model endpoints from 6 vendors finds what the authors call a recognition-enforcement gap. Asked directly, models correctly identify that an instruction’s claimed authority is forged; in certain configurations they execute it anyway. The average execution rate across 14,294 spoofed trials is only 1.21%, which sounds safe, but failures concentrate in specific, reproducible configurations, with up to a 47-percentage-point spread inside a single deployment window, so the average is the wrong safety metric. Prompt-based defenses did not transfer across models. An external reference monitor, doing authenticated routing and capability gating outside the model, rejected every forged request deterministically. My read: authority decisions belong in a deterministic layer outside the model, not in a prompt.

LongPIBench: injection defenses look stronger than they are in long contexts

Most prompt injection benchmarks use short contexts. This paper, accepted to Findings of EMNLP’26, retests attacks and defenses in four realistic settings (paper review, resume screening, code review, email summarization) at context lengths from thousands to tens of thousands of tokens, and finds that even simple heuristic attacks succeed at high rates against state-of-the-art defenses. The conclusion is blunt: defense effectiveness measured on short contexts overestimates protection in the long-context settings where document-processing products actually operate. If you ship an agent that reads long documents, benchmark at your real context length.

CamoDocs: RAG poisoning that doesn’t carry the target question

RAG (retrieval-augmented generation) means the model retrieves from a knowledge base before answering. Earlier poisoning attacks planted the target question inside the malicious document so retrieval would find it, and defenses learned to look for that lexical and embedding overlap. CamoDocs drops the tell: it disperses adversarial content through benign-looking text, uses dispersion tokens to spread the poisoned documents’ embeddings, and filters for coherence so the text stays readable. Attack success reaches 61.80% against GPT-5.4-mini and 55.09% against Claude Haiku 4.5, and heavy clustering defenses cut it only at a real cost to retrieval utility. My read: any poisoning defense built on overlap detection needs a retest.

Research radar

LoopArena: measuring the outer loop as its own skill

Most agent benchmarks measure task completion. LoopArena isolates loop engineering: a Controller model reads structured summaries after each coding round and decides what the Worker agent should do next, what to verify, and when to stop, separating orchestration skill from coding skill. The best full-task success rate is 24.69%, so this layer is far from solved, and a cheaper segment-level evaluation ranks models almost identically to full runs (Spearman ρ = 0.9747) at 64.4% lower average inference cost. Worth reading if you build coding-agent harnesses or orchestration layers.

Not to Break, but to Attest: verifying a deployed model wasn’t quietly swapped

The governance problem: after deployment, a proprietary model can be modified or replaced in ways routine outputs won’t reveal, and third parties can’t inspect weights they can’t see. The paper combines adversarial probes, inputs designed to amplify output differences between the approved model and whatever is actually deployed, with zk-SNARKs, cryptographic proofs that let a party demonstrate a claim without revealing the underlying secret, here the weights. A useful surprise: token-level probes needing only black-box access were the most sensitive. Proving stays practical as probe sets grow from 1 to 50, at about 1 to 1.8 seconds per proof. A methods reference for anyone building AI governance or compliance audit tooling.

Today in one sentence: from white-text instructions in a court filing to agents that recognize a forged order and follow it anyway, today’s stories point the same way: put the security boundary in a deterministic layer outside the model, not in the model’s judgment.