Can researchers trust OpenAI with unpublished math?
Andreas Thom, a group theorist at TU Dresden, posted an account on Mathstodon: he spent months discussing problems around sofic groups with ChatGPT. Then OpenAI announced the construction of a non-sofic group, settling a question Mikhail Gromov posed 27 years ago, with a proof that by Thom’s account builds directly on his 2019 paper with Gábor Kun. He has written to Mark Sellke and to OpenAI’s Sébastien Bubeck with two distinct questions: were his conversations used in training, and did the system read them while working on the proof. The only reply so far, which Thom has made public, is a single line from Sellke — “that did not happen” — which leaves unclear how much of either question it covers; OpenAI has made no formal public response. For the field, though, the doubt itself is the story: a researcher who cannot rule this out will keep their best unpublished ideas out of the chat box. Consumer ChatGPT conversations can feed training unless you opt out; zero-retention terms are written for enterprises, not for a professor with a conjecture. After this week’s Navier-Stokes credit dispute, this is the same trust crack widening.
OpenAI ships the Agents API: the Codex harness becomes a service
On September 10 OpenAI opened a public beta of the Agents API. A harness is the machinery wrapped around a model: it feeds context, calls tools, retries after failures. Until now that layer was typically something teams built themselves; the Agents API rents out the managed one that runs Codex. You supply instructions, tools and MCP servers; OpenAI keeps the session alive, compacts context when it fills up, and recovers from crashes. Sessions can run code, edit files, search the web, produce artifacts and fan work out to subagents under a concurrency cap, billed for the tokens and tools the agents use rather than a separate API fee. My read: the agent-infrastructure race is shifting from whose model is stronger to who manages your state. The price is equally clear. The orchestration — session state, compaction, recovery — now lives with OpenAI even when execution runs in your own environment, so switching vendors means swapping a runtime, not an endpoint.
Anthropic’s threat intelligence report: eight months, seven abuse categories
The report covers abuse disrupted between December 2025 and August 2026, tracked across seven categories: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons, and distillation (training your own model on another model’s outputs to copy its capability). Two case details stand out. A fraud ring used commercial CAPTCHA-solving services to mass-register accounts on exchanges. An influence network built roughly 1,000 social accounts with warm-up behavior so they would look human before deployment, then paired fake outlets with journalist bylines and AI-generated headshots across continents. The report also notes that only one case, an illicit distillation, involved Fable- or Mythos-class models. My read: the barrier for abusers is process engineering, not model capability. CAPTCHA-solving and account warming are off-the-shelf gray-market services; AI fills in the content step. And distillation now sits beside bioweapons and cyber in a frontier lab’s abuse taxonomy, which tells you labs treat capability theft as a security problem rather than a commercial dispute.
Cognition’s SWE-2: Kimi K3 post-trained to the frontier tier
SWE-2 is a coding model built by reinforcement-learning post-training on Kimi K3, Moonshot AI’s open-weight 2.8-trillion-parameter mixture-of-experts model. By Cognition’s own benchmarks it scores 50.0% on FrontierCode 1.1 Main, within one point of Anthropic’s Fable 5.1, at 64% lower cost; it is available today in Devin Desktop and CLI. The signal is the recipe, not the score: an American agent company takes a Chinese open-weight base, adds its own RL, and lands within a point of a closed frontier model for a fraction of the price. Open weights are becoming the shared substrate of the agent industry, and closed labs will need something harder to copy than post-training to defend their lead.
The rumor is the exploit
Anil Madhavapeddy, Cambridge professor and long-time OCaml contributor, reports that knowing only roughly what a security vulnerability was about, his own AI agents independently rediscovered it and produced a working exploit before the public patch shipped. That breaks a load-bearing assumption of open-source security. Coordinated disclosure (keeping vulnerability details secret until a patch is ready) rests on the bet that vague hints alone don’t hand attackers a working exploit; in his test, a rumor plus an agent produced one in under a minute. This is one practitioner’s test, not a systematic study, but the assumption it breaks is one every maintainer relies on: patch and disclosure now have to land together, because the embargo window itself no longer protects much.
Watermarks on, verification missing
Since August 2, Article 50 of the EU AI Act requires machine-readable marking of generative AI output. Anthropic embeds a SynthID-Text variant in new Claude models at the model level, worldwide, with no user off-switch (see the official explainer); Google’s Gemini has carried SynthID-Text since 2024. The watermark changes only the randomness used to pick each next word: invisible to readers, statistically readable by detection tools. The paper’s argument is that nobody outside the vendors can currently check either side of the debate, neither the side effects critics worry about nor the detection rates vendors promise, because there are no matched output samples, no configuration disclosure, no accredited audits, and no interoperable detectors. The watermarks are shipping; the means to verify them are not. The paper calls that gap the real governance failure, and I agree: this is exactly where content-labeling rules stall on paper.
$0 license fees: OpenAI and GSA extend the government deal to every level
OpenAI and the General Services Administration, the US government’s central purchasing agency, signed a 27-month agreement running October 2026 through December 2028: federal, state, local and tribal governments get ChatGPT with the standard $15-per-user monthly license fee waived and usage at half price, plus expanded cyber-defense support. Roughly 23 million public-sector employees become eligible, with state and local governments included for the first time. The earlier $1-per-agency-per-year arrangement becomes a metered long-term contract (see Nextgov). Pricing has moved from customer acquisition to market cultivation, and for government buyers the real cost shifts to usage: budget offices now have to learn to forecast token consumption.
Astra demand pauses $200 Pro sign-ups
One week after GPT-6 Astra shipped on September 3, OpenAI’s VP of Product and Platforms Thibault Sottiaux announced that new sign-ups for the $200-a-month Pro plan are paused: “This tier puts the heaviest load on our systems.” No timeline was given; other tiers and the pay-as-you-go API are unaffected (announced only on X, no press page). Compute supply is the binding constraint again. If your production setup depends on one vendor’s top tier, this is your reminder to keep a second path.
Research radar
How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE
Directional ablation is a white-box attack: a few hundred contrastive prompts locate the single activation-space direction a model uses to encode refusal, which is then projected out of the weights, no training required. Previously shown only on small and mid-size dense models, here it is scaled to GLM-5.3-Flash, a 320-billion-parameter mixture-of-experts model. The naive recipe fails silently on MoE, but with adapted targeting refusal drops by 41 to 89 percentage points across seven harm benchmarks. If you work on open-weight release decisions or tamper-resistant alignment, read it: scale did not buy sturdier alignment.
Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning
The standing assumption was that cipher-based jailbreaks need a fine-tuning API to teach a model to converse in code. This paper gets there with prompting and in-context learning alone, jailbreaking frontier models from Anthropic, Google and OpenAI; because the harmful content travels encrypted, output-side harm classifiers miss it too. Relevant if you build safety classifiers or run red teams: locking down fine-tuning no longer closes this attack line.
Show-Harness: Just a VLM Agent Can Play Robots
A compact interface of discrete semantic action units turns robot control into the kind of semantic reasoning a general-purpose vision-language model already does: closed frontier VLMs drive robots zero-shot, small open-source VLMs need a few GPU-hours of fine-tuning, and the resulting agents generalize across tasks and robot bodies better than representative VLA systems (VLA: vision-language-action models that train control directly into the weights). Worth a click for embodied-AI researchers: it questions the assumption that control has to live in the weights.
One line for today: models keep proving what they can do, while the mechanisms that would let anyone verify what they didn’t do (read your chats, miss with their watermark, leak an exploit window) still don’t exist. Today’s top story, the watermark paper and the exploit post are all missing the same piece.