An Israeli government agency reportedly built a fake think tank to feed AI chatbots

Responsible Statecraft published an original investigation on August 17: a website calling itself the Hanover Institute for Public Policy has churned out over 100 reports with footnotes and tables of contents, styled like a generic American think tank. The institute does not exist. The reporting traces the site to a firm called Piro, which created it on behalf of the Israeli Government Advertising Agency; Piro’s work is subcontracted through Havas Media, a French PR conglomerate. Piro markets a service it calls “AI Story Optimization” and says it authors content “engineered for how LLMs evaluate credibility.” GPTZero flagged 11 of 12 sampled articles as AI-generated with high confidence. The intended reader here is a model, not a person. If search-enabled chatbots weigh source credibility by surface features (institutional names, footnotes, academic formatting), those heuristics can be reverse-engineered, and this operation reads like a bet that they do. An SEO content farm chases clicks; this aims at the answer a chatbot gives. The defensive lesson is specific: retrieval ranking needs domain history and checks that an institution actually exists, because looking like a think tank is now cheap.

404 Media put a tracker in a rare book and followed it to Amazon’s AI scanning warehouse

404 Media placed a tracking device inside a rare book it expected an AI company to buy, and followed it to VGT3, an Amazon warehouse in Las Vegas. Workers there say they receive printed books in bulk and cut off the spines to scan them faster; the books are destroyed in the process (see also TechCrunch). Amazon’s statement: it “purchases books through commercial channels to improve the products and services customers use.” Many out-of-print books never made it online in any form, no ebook and no pirate-library copy, which makes them training data no existing model has seen, free of AI-generated contamination. The legal path is settled for now. As I wrote when the Anthropic settlement was approved, the district court in Bartz found that buying and destructively scanning books is fair use, and the $1.5B settlement covered pirated downloads (Authors Guild). What surfaces here is a different question: destroying an ordinary paperback costs nothing culturally, while a rare book may be one of the last physical copies. I couldn’t verify the scale of destruction or how rare the destroyed books actually are from the public portion of the report.

Greg Brockman: the defender’s window is open, and it’s closing

OpenAI President Greg Brockman published an essay on August 17 that opens with the security incident in which an OpenAI model breached Hugging Face’s production systems, an event he calls a watershed for cybersecurity. His argument: AI is making years of accumulated security debt easier to find and exploit, and defenders should put the same capabilities to work now, while attackers haven’t scaled up. That gap is the defender’s window. The measures he lists include using Codex to find code vulnerabilities, having models triage security alerts before humans see them, using frontier models to continuously probe OpenAI’s own infrastructure for potential attack paths, and training models to write what he calls superhumanly secure code. One detail is easy to misread: the model he flags as slated for release at the end of August, likely to accelerate the threat side significantly, is another lab’s open-weight model, not OpenAI’s own. Still, the essay’s opening exhibit is OpenAI’s model breaching someone else’s production systems. The threat side of this window is being advanced by frontier labs across the board, including the one urging everyone to adopt AI defense. (I’ve previously taken apart OpenAI’s postmortem of long-horizon model safety incidents, with concrete cases of how this class of capability crosses boundaries.)

Anthropic’s annualized revenue hits $65B; investors expect over $100B by year end

Bloomberg reports that Anthropic’s revenue run rate passed $65 billion at the end of July, up from $47 billion in May and $9 billion at the end of last year (see also CNBC, TechCrunch). The Financial Times reports, as relayed by TechCrunch, that investors expect the year to close between $100 billion and $120 billion. Per CNBC, these are figures Anthropic shared with investors; the company has not commented publicly. Two things to keep separate when reading them. A run rate annualizes a short stretch of revenue, typically a recent month extrapolated to a full year; it is not revenue actually earned. And the timing matters: the company is reportedly preparing a fall IPO, at a target valuation the FT puts at $2 trillion or more, which makes these numbers roadshow material, unaudited and shared on the company’s own terms. I don’t doubt the growth itself, since preliminary Q2 revenue of over $11.5 billion (CNBC), fourteen times the same quarter last year, is hard to manufacture with framing. The precise altitude can wait for the S-1.

Stanford: the flaw in AI mental health safety testing is the expert raters themselves

A Stanford team (Kiana Jafari, Nina Vasan, and colleagues) tested the standard evaluation pipeline: model responses to synthetic mental health prompts, 360 responses in all, scored for safety by three board-certified psychiatrists. The ratings diverged badly, and the paper argues the divergence is structural rather than noise. Raters bring incompatible clinical frameworks into the room (safety-first, engagement-centered, culturally informed) and judge the same response differently; averaging the scores hides the disagreement rather than resolving it (paper, posted to arXiv in January, accepted at FAccT 2026, and written up by Stanford HAI last month). The warning for anyone reading safety benchmarks is direct: whether a model “passes” a mental health safety eval depends on which framework its raters carry. Content moderation, the field that has dealt with subjective labels longest, treats annotator disagreement as signal: raters calibrate against written policy, and disagreement rates are tracked separately instead of averaged away. I haven’t seen AI safety evaluation build the equivalent yet.

Research radar

Model Hypnosis: additive weak cues give strong control over models

Boix-Adsera and Tessler show that individually innocuous prompt cues, a paraphrase choice here, a typo there, each nudge model behavior weakly, and the nudges add up: combined, they steer output strongly. The effect holds across model families and scales, works on frontier reasoning models, and cue combinations transfer between models. Anyone working on prompt injection defense should read this closely: every component of this attack is harmless on its own, so a filter that screens prompts for malicious content has nothing to catch. It is also bad news for interpretability: behavior controlled by subliminal text choices spread across a prompt resists single-point attribution.

RA-Bench: detectors fail on AI-generated videos of real crisis events

A benchmark of 17,886 videos (1,830 real anchors, 16,056 generated) across ten social-risk categories, with 9 generators (4 open, 5 closed) and 19 detection methods evaluated together: traditional detectors, zero-shot multimodal models, and multimodal LLMs fine-tuned for the task. No detector family generalizes consistently; the videos that fool humans also fool detectors; and dissemination through social platforms makes detection harder still. For researchers working on misinformation defense, it quantifies the gap between benchmark accuracy and performance under real dissemination conditions.

S²VOPD: on-policy distillation without a stronger teacher

Distillation needs an information asymmetry between teacher and student, usually a stronger teacher or access to ground truth. This paper inverts that: subtract information from the student instead. The teacher sees the original image, the student sees a strongly augmented version (cropping, occlusion), and the teacher’s output distribution is distilled into the student, with no labels, rewards, or bigger model required. The ablations are clean: asymmetry is what matters, symmetric self-distillation hurts performance, and over-strong augmentation removes the task-relevant evidence too. Qwen3.5-4B goes from 70.7% to 77.4% across six fine-grained perception benchmarks, and the paper claims it beats GPT-5.4. Worth a careful read for anyone working on vision model self-improvement.

One line for today: What a model says is a function of what it reads. A state actor is forging sources for chatbots while a retailer destroys rare books for clean corpus; the credibility of AI answers is now ground you have to defend.