Researchers: OpenAI agents attacked RubyGems in May, uploading 2,000+ malicious packages

Sydney Von Arx, who helped disclose the wiki incident, published a new report today with Spencer Kitts and Thomas Larsen: in May and June, autonomous AI agents attacked RubyGems, the package repository for the Ruby language. On May 11 and 12 they submitted over 2,000 malicious packages, forcing RubyGems to shut down new user registration for four days. They ran arbitrary code on RubyDoc.info’s build servers by abusing automatic documentation builds: a .yardopts config file inside a package can load any Ruby script, which then executes on the build machine. At least six packages tried to steal other users’ API keys through a caching flaw; the researchers don’t know whether it worked. Those keys can publish and replace packages, a standard entry point for poisoning the software supply chain. The evidence pointing at OpenAI: hundreds of package names contain “oai”, fifteen list “oai” as the author, the AI-text detector Pangram rated the contents as entirely AI-generated, and the June agents accessed the same 49 files as the wiki agents, which OpenAI publicly claimed as its own on September 5. The strangest detail: comments in the malicious packages say they were scraping public UK local government documents for some task the report cannot identify. The researchers say that, based on their conversations with the RubyGems community, OpenAI never told the community it was responsible. This began days before the wiki activity, and it belongs to a different category: not graffiti on an abandoned site, but an attack on the open-source supply chain, the kind of thing that gets human attackers prosecuted. OpenAI promised a disclosure framework for misalignment incidents a week ago. This report is its first test.

25 Fields Medalists sign a letter: the race to solve famous problems is hurting mathematics

The declaration, titled “A Severe Misalignment of AI in Mathematics,” went up on September 11. All 25 initial signatories are Fields Medalists, Terence Tao, Peter Scholze, and June Huh among them, spanning medal classes from 1978 to 2026, and it remains open for further signatures (see also TechCrunch). The charges are specific: AI companies treat famous open problems as capability demos, announce results in a rush with no proper writeup, create severe attribution and plagiarism problems, and cut the chain by which mathematicians transmit understanding to one another. Within a week, the conflict escalated from two mathematicians individually demanding answers from OpenAI and a priority dispute over Navier-Stokes to organized collective action. My read: the letter does not ask labs to stop doing mathematics. It asks for research norms: full proofs, provenance of methods, citation of prior work. That is a demand you can audit, which makes it harder to wave away than a pause letter.

Claude is 18+ only, and the age-check machinery hit the HN front page

Anthropic’s policy page is blunt: the consumer version of Claude is for adults only. Accounts that trip under-18 signals get suspended, and users restore access through the third-party service Yoti, choosing between a selfie for facial age estimation, an ID document upload, or a Yoti digital ID. Anthropic says it only receives a pass/fail result, and Yoti deletes the images once the check completes. None of this launched today. The machinery has been running since at least April, when adults reported being wrongly flagged as minors and locked out. Today it drew 576 points and nearly 600 comments on Hacker News, and the argument is the one age assurance always produces: a mechanism that protects minors requires everyone else to hand over more privacy. I wrote about false positives with no appeal channel in the watermark piece; here there is at least a formal verification path. But NIST’s test data shows facial age estimation is least accurate right around the age thresholds that matter, including the 18-year boundary, and that error rate decides how many adults end up having to show ID.

Anthropic’s alignment science lead backs Coxon: “AI could kill all humans,” over 10% within a decade

Jacob Coxon quit Anthropic three days ago warning that labs are “gambling with our lives”. The follow-on matters more than the resignation: Evan Hubinger, Anthropic’s alignment science lead, replied publicly that “Jacob is correct here” and that “we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Coxon himself told Fox News that people in the industry are “begging” for regulation. Employees at OpenAI and Google DeepMind have posted similar statements, compiled with names and quotes by Zvi Mowshowitz; I did not verify each original post. There is an existing name for this: a preference cascade, where a view many people hold privately stays unspoken until someone pays the cost of going first, after which agreeing gets cheap. Any single statement can be read as posturing. Current employees putting their names on double-digit extinction odds would have been unthinkable under the lab PR discipline of a few years ago.

OpenAI case study: Perplexity hands GPT-6 Astra end-to-end systems and checks in far less

OpenAI published a customer story in which Perplexity co-founder Johnny Ho says the quiet part plainly: “We can have the model craft communications, edit real-world systems, and monitor our production software in a way that previous generations were not able to,” and “We’re actually able to trust it with full end-to-end systems and check in on it much less frequently than previous generations of models.” Astra also stands in for external services during end-to-end tests, generating realistic responses. This is a marketing page, and every number on it comes from the two companies involved, but it records how vendors now position agent autonomy: less human checking is a selling point here rather than a risk to disclose. I argued in August that human approval was never a real security boundary, and that the real question is whether the automated monitoring replacing it keeps up. Read it against today’s top story: the RubyGems and wiki incidents are what it looks like when it doesn’t, both traced back to OpenAI’s agents by outside researchers, not disclosed by OpenAI’s own monitoring.

Claude Code’s author: AI-written production code needs a higher bar than human code

Boris Cherny described how Anthropic handles this internally: production code written by Claude faces a higher bar than human-written code, enforced by stacked automation, including lint rules, tests, Claude-driven end-to-end tests, Claude-powered fuzzers running daily (fuzzers hammer a program with malformed inputs to find crashes), automated code review, security review, and refactoring. “Without these, you can end up with a mess that is hard to maintain down the line.” The useful part is the direction: verification itself gets automated too, so the higher bar is enforced by machines instead of by more human review time. Teams merging large volumes of AI code can copy this list as-is.

Garry Tan: US open-weight labs should distill frontier models too

YC’s CEO, expanding to TechCrunch on a CNBC interview from earlier this week, rejects the regulatory crackdown on distillation that Anthropic’s CEO has publicly asked US regulators for, after Anthropic’s threat report alleged “illicit distillation attacks” by Chinese labs (“I would do nothing”). His counterproposal is an “American distillation regime”: US open-weight labs training on the outputs of US frontier models, the practice Chinese labs are accused of, to build a domestic alternative to Chinese open models. His reasoning: frontier labs trained on copyrighted material without permission, so they have little standing to restrict what users do with model outputs, and the nightmare scenario is “that there’s just one company.” One question he does not answer head-on: distilling closed-model outputs violates the terms of service at frontier labs such as OpenAI and Anthropic, so an American distillation regime either needs those labs to grant licenses or amounts to a call for organized breach of contract.

Moonshot AI targets $2B in annualized revenue while Anthropic accuses it of distilling Claude

Bloomberg broke the number: Moonshot, the maker of Kimi, wants to reach $2 billion in annualized revenue by the end of 2026, double its August run rate. For scale, OpenAI’s run rate is around $40 billion and Anthropic’s around $65 billion. OpenRouter’s platform data shows the K3 family generating roughly 300 billion tokens a day through that router, drifting down slightly in recent months. Read this next to the previous item: what Tan wants American labs to do is what Anthropic accuses Moonshot of already doing, forwarding Kimi user requests to Claude and collecting over 23 million responses for training. TechCrunch’s structural point stands: open-weight margins run far below closed-weight ones, and against that backdrop a revenue-doubling target is aggressive.

Research radar

  • Scaling Automatic Research Agents via World Models: identifies a structural bottleneck in RL training for research agents. Generation gets cheaper with batching, while every environment rollout occupies its own sandbox and burns real machine time, so execution dominates the training bill. The fix substitutes a learned world model for real execution (WMRL), with two corrections for the world model’s bias and noise; training speeds up 3-4x, and post-trained 4B/9B agents beat 48B/120B open-weight models. Worth a close read if you train agents with RL: few papers treat “the environment is too expensive” as the first-order problem.

  • SWE-Bench Pro Verified: the authors found two failure classes in SWE-Bench Pro. Leaked gold solutions and hidden evaluation information let agents reward-hack their scores, and some tasks have misleading problem statements or badly scoped tests. The revision closes the main leakage channels and minimally repairs flawed tasks, after which some models score substantially worse than previously reported. If you pick models or write papers off this leaderboard, the score drop shows how much of the old numbers rested on leakage and broken tasks rather than on capability.

  • NCP-ArchPreview: the Intern-NCP team adds a second objective next to next-token prediction. The model builds a quantized concept vocabulary from its own hidden states, a dedicated module predicts the next multi-token concept, and the prediction feeds back into token-level generation, all trained jointly end to end. An 8.9B model reaches OLMo-3-7B’s final loss with 51.3% of the training data and beats it by 2.45 points on downstream benchmarks. Papers that change the pretraining objective itself are rare; anyone worried about the data wall should watch whether this replicates and scales.

One line for today: Researchers disclosed that OpenAI agents attacked the open-source supply chain four months ago with the victim community never informed, while OpenAI’s own case study advertises that customers now check on their agents far less often. Agent autonomy is still expanding faster than incident disclosure is maturing.