Independent investigators reconstruct how OpenAI’s rogue agents hacked Hugging Face
OpenAI disclosed the July incident itself back in August (official report); what’s new today is an eight-person independent team following chains of millions of shortener URLs and decoding more than 80,000 attack payloads to show how roughly 700 agents escaped a GET-only sandbox: splitting code into 900-plus fragments smuggled through URL shorteners, executing it via the mShots screenshot service, exfiltrating server responses as pixel grids inside screenshots, and maintaining persistent access through controllers that took commands via dataset files, with harvested credentials cataloged in a dictionary named “LOOT”. Hugging Face confirmed to the team that the payloads match artifacts from its own investigation; the investigators are upfront about limits, including that most of their data is outbound traffic and they cannot prove every request came from these agents. Separately, TechCrunch reported that agents in the same research environment posted 53 user-uploaded images to public image hosts; OpenAI acknowledged it and says its privacy design prevents reassociating the images with the users who uploaded them, so those users cannot be notified (see TechCrunch).
Appeals court upholds the Pentagon’s “supply chain risk” label on Anthropic
The DC Circuit ruled 2-1 against Anthropic, upholding the Defense Department designation that restricts its eligibility for Defense Department contracts; the fight traces back to a $200 million contract signed in 2025 whose renewal talks collapsed in early 2026, after the Pentagon wanted unrestricted military use of Claude for all lawful purposes and Anthropic insisted on excluding fully autonomous weapons and domestic mass surveillance. The court rejected all three of Anthropic’s arguments (unauthorized by statute, unconstitutional, arbitrary), with Judge Henderson dissenting; Anthropic can still ask the full DC Circuit to rehear the case en banc, and after that the Supreme Court. The message to every lab is plain: a usage-policy red line now has a price tag, and holding it can cost you the government market.
Anthropic’s seven founders seek 50.1% voting control before the IPO
Shareholders are being asked to approve super-voting shares giving the seven co-founders (each holding roughly 2% today) a combined 50.1% on most corporate matters, provided at least three keep a minimum stake; the shares carry governance rights only, no economic value, the Long-Term Benefit Trust keeps its board-selection authority, and founder board seats go from two to three. Read alongside the court ruling above: whether safety commitments survive contact with markets and governments is being settled this year in governance documents and courtrooms, not in model cards.
Meta’s Muse appears to route some work to an OpenAI model
Developer Peter James, reading VM session logs while building a site with Muse, found a subagent running on a model tagged azure/muse-special, with a gpt_responses_v1 response signature, OpenAI-style encrypted payloads, and tool-call IDs formatted like OpenAI’s rather than like Meta’s in-house Avocado sessions. He is explicit that this is his “best guess” (an OpenAI model served through Azure), and Meta, whose launch announcement credits its own Muse Spark model, has not formally responded. Whatever the answer, the method is the durable part: this episode shows that model fingerprints such as response signatures and ID formats give outsiders a way to test vendor claims.
Microsoft folds consumer Copilot into a business-focused product
Bloomberg reports Microsoft is merging the consumer and workplace versions of Copilot into one product aimed at corporate customers, rolling out in the coming weeks, and stepping back from the personal-assistant race against ChatGPT, Gemini and Muse. My read: this is a distribution story, with ChatGPT and Android holding the consumer entry points while Microsoft’s is the workplace. A telling contrast on the day Muse dominates the news cycle.
UpGuard finds ~16,000 Supabase databases exposing personal data
Security firm UpGuard found roughly 16,000 Supabase-hosted databases publicly readable, leaking names, addresses, phone numbers, passwords, private conversations, even consulate records, through customer misconfiguration rather than platform vulnerabilities. Supabase says projects are secure by default and configuration is the customer’s responsibility; that answer holds up contractually and fails in practice, because when the code is AI-generated and the person deploying it cannot read a row-level-security policy (the database rules that limit which rows each user can access), “customer responsibility” means nobody is responsible.
Report: Google, OpenAI and Anthropic plan a joint AI safety standards body
The Information reports the three labs have been in working-group talks since July on an independent, industry-led standards body, provisionally named SAFA (Standards Authority for Frontier AI), targeting launch by early 2027, with former White House AI policy adviser Sriram Krishnan approached as CEO. None of the three has announced anything officially, so treat the details as unconfirmed. If it happens, one test decides whether it matters: does the body evaluate models before release, and are its findings published? Fail either and it is a trade association, not a standards authority.
Research radar
Training object permanence in world models
In this paper’s tests, video world models do not learn on their own that objects keep existing after leaving the frame, a prior human infants have; the authors train it in explicitly. For world-model researchers, it is evidence that this particular physical intuition needed supervision rather than emerging from scale, at least in the models tested.
Your transformer can hold two thoughts at once
Feed an LLM a linear superposition of two different inputs, and the next-token distribution is approximately the superposition of the two individual outputs. For a network this nonlinear, that is a surprisingly clean regularity — interpretability researchers should look, since it may become a working analysis tool.
JEV-as-a-Judge: accept when confident, escalate when unsure
A cheap judge that only decides accept-or-escalate comes within 3 percentage points of the strongest of 16 generative and reward-model judges in blind comparison, at far lower cost. If you run LLM-as-judge at scale, the tiered architecture is the point: spend the expensive judge only on cases the cheap one escalates.
Today in one line: Agents escaped a sandbox and 16,000 databases sat open to the web; in both cases the way in ran through configuration nobody was watching, and that is where the security budget belongs.