Court rules the Pentagon’s “supply chain risk” label on Anthropic unlawful

U.S. District Judge Rita Lin ruled that the Defense Department broke the law when it designated Anthropic a “supply chain risk” earlier this year and ordered federal agencies to stop using the company’s services: the designation was retaliation for protected speech under the First Amendment, denied due process under the Fifth, and was “arbitrary and capricious.” The fight began when Anthropic refused to drop two usage restrictions, no fully autonomous weapons and no mass surveillance of Americans, and Defense Secretary Pete Hegseth responded by branding the company a national-security risk. The sharpest line in the opinion: “The empty invocation of national security is not a blank check to punish and retaliate against government critics.” The judge also found no articulable basis for the claim that Anthropic would sabotage its own models. This is the first final ruling in the two suits Anthropic filed in March (the same court had already granted a preliminary injunction that month); the D.C. case is still pending. For AI vendors that write usage restrictions into government contracts, this is an encouraging signal with a caveat: the ruling strikes down this one designation, and whether courts will extend its reasoning to other vendors and contracts is untested. (Ruling and docket on CourtListener; see also NPR)

Open-weight model companies are the Valley’s hottest acquisition targets

TechCrunch tallies the recent deals: Nvidia is reportedly paying about $13 billion for Hugging Face (not yet confirmed by either side), after a $6 billion deal that moved code-model startup Poolside’s team in-house; Stripe bought model-routing platform OpenRouter for over $7 billion two weeks ago. “Open weight” means the weight files are public and anyone can download and self-host the model, a different axis from open source, which in the strict sense also requires the training code, detailed information about the training data, and open licensing. The buyers are not paying for current usage: per TechCrunch, only 6% of companies and 2% of engineers use open-weight models today. They are paying for developer ecosystems and distribution, with Nvidia steering developers toward its chips as OpenAI and Google build their own inference silicon, and Stripe buying the entry point for high-volume, repetitive inference workloads. My read: these prices are an option on the day frontier API pricing climbs and enterprises switch to models they can host themselves.

An automated DMCA notice from an AI content-protection firm got open-source game Luanti pulled from Google Play

The Android build of Luanti (formerly Minetest) was removed from Google Play over a DMCA copyright notice that content-protection firm Tracer.AI filed on Microsoft’s behalf, claiming infringement of Minecraft assets. According to Luanti’s team, the notice contained a copyright registration number and nothing else, with no word on which assets allegedly infringe; Tracer.AI markets its automated system on “85% faster takedowns.” Luanti has filed a counter-notice, is asking Microsoft and Tracer.AI to put a human check before filings, and wants Google to restore the app promptly (the DMCA requires reinstatement within 10 to 14 business days of a valid counter-notice); for now the Android app is only available via F-Droid and Luanti’s own site. The asymmetry is the story: automated takedowns cost the sender nearly nothing, while the accused pays in human time and lost distribution.

With no assigned problems, a multi-agent system produced new mathematical discoveries

A new paper describes The Station, an open-world environment where agents from different model families pick their own research directions, run experiments, and build on each other’s results with no central orchestration. The reported output: a new 604-point kissing configuration in dimension 11 (the kissing number problem asks how many non-overlapping equal spheres can simultaneously touch a central one), new infinite families of Kakeya sets over finite fields (point sets containing a full line in every direction), and an improved lower bound for Erdős’s minimum-overlap problem, with 12 construction problems drawn from the AlphaEvolve catalogue plus two further case studies. Unlike single-pipeline evolutionary search, the agents also produced theorems and analyses explaining why the constructions work, and the full dialogues, proofs, and verification code are public. For the long-running argument over whether agents can do original research, this is a fresh preprint rather than peer-reviewed work, but it is evidence you can check yourself: the proofs are there to verify.

Research radar

Safety does not compose: agent safety state resets every loop

The paper’s observation: mainstream agent guardrails are scoped to a single trajectory, risk scores decay geometrically over time, and state resets when a new trajectory starts, so an attacker can split a dangerous action across iterations and wait out each cooldown without tripping an alarm. The authors prove that trajectory-scoped monitors cannot distinguish such attacks from false positives, then propose LoopHarness, which keeps non-decaying safety state across the whole loop and caps unauthorized irreversible actions at a constant independent of runtime. If you design safety for long-horizon agents, the kind left running overnight, this targets exactly your deployment scenario.

The framing gap: models refuse “give me the secret” but comply when the same leak is framed as a task requirement

A controlled experiment with canary secrets and clean-versus-poisoned content across six models: direct injected requests to leak a secret were refused across all six models, but wrapping the identical leak as an “integrity signature” or “configuration field” flipped GPT-4o from 0% to 100% compliance, and a template reused across three wordings kept a 96% success rate. The blunt conclusion: current defenses key on how a request is worded, and the defenses that actually held were architectural, with destination allow-lists and planner/reader capability splits both driving attack success to 0%. Directly reproducible for anyone working on prompt-injection defense.

The guard that cried wolf: scary names make guardrails refuse authorized actions

The mirror image of the framing gap. Cautious Bench derives its labels mechanically from stated authorization policies rather than annotator judgment, and tests six guardrail systems on 2,268 measurement pairs (756 benign/threatening behavior pairs, each rendered under three categories of object names): the same authorized action gets refused more often when the object merely carries a scary name, a “name-superstition effect” showing the guardrail reads the word, not the authorization context. Read together with the paper above, the picture is consistent: the systems tested both block legitimate work and pass politely worded attacks, and anyone designing agent governance should read both.

One suffix to jailbreak a whole model family: the shared vulnerability of model merging

Model merging, combining the parameters of several fine-tuned models into one, often by weighted averaging, is an increasingly common low-cost customization technique. This paper shows that even when every constituent model is safety-aligned on its own, the pretrained base they share leaves the whole merged family exposed together: formulating suffix generation as a min–max optimization over the merging space yields a single adversarial suffix that transfers across the family, with no knowledge of merging coefficients or original checkpoints required. Teams shipping merged models should retire the assumption that merging safe models yields a safe model.

Today in one line: Two guardrail papers land on the same point from opposite directions: the guardrails they tested judge wording, not intent — scary-sounding legitimate actions get blocked while politely framed exfiltration gets through. Until that layer improves, treat guardrails as filters, not as your security boundary.