Paul Christiano joins the OpenAI Foundation board
OpenAI announced that alignment researcher Paul Christiano is joining the OpenAI Foundation board, where he will sit on the Safety and Security Committee chaired by Zico Kolter and also attend the board of the for-profit OpenAI Group PBC as a non-voting observer. Christiano led OpenAI’s alignment research from 2017 to 2021 and did foundational work on RLHF (reinforcement learning from human feedback); he later founded the nonprofit Alignment Research Center and is currently Senior Technical Advisor to the Center for AI Standards and Innovation at NIST, the US standards agency, where he will now recuse himself from OpenAI-related matters and model evaluations. My read: seating a researcher known for taking extinction risk seriously on a safety governance body is one of OpenAI’s weightiest governance moves in years. What the announcement doesn’t say is whether the committee can veto or delay a launch, and that is what the seat is actually worth. The same day, OpenAI published a call to move fast on safety standards while the policy window is open.
Anthropic researcher quits with a warning about self-improving AI
Jacob Coxon, a 27-year-old researcher who spent the past three years on pretraining at OpenAI and Anthropic, resigned from Anthropic with a seven-part thread on X that, per Deadline, drew nearly 76 million views overnight. His words: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.” He calls for pacing agreements between US labs, meaning coordinated limits on how fast capabilities advance, possibly including a temporary freeze on capability improvements; he argues recent incidents such as the attack on Hugging Face make such agreements more feasible (I have not independently verified the details of that incident). Read together with the item above: on the same day, one alignment researcher chose a board seat and another chose the exit. Both choices rest on the same judgment, that the labs’ self-restraint can no longer be taken on trust.
WeWorm: a zero-click worm that spreads through WeChat calls
Calif Research published WeWorm, a proof-of-concept worm built on a memory-corruption bug in WeChat’s VoIP stack. An attacker places a WeChat call; the exploit fires while the phone is still ringing, before the victim answers or touches anything, and hands over full control of the WeChat account: reading and sending messages, placing calls, and automatically dialing the next victim, on both iOS and Android. The number worth remembering is the timeline: with AI assistance, the team had a working Android remote code execution (RCE, meaning the attacker can run arbitrary code on the target device) proof of concept in about two days, and a full worm demo after one more week. Disclosure was orderly: reported to Tencent on July 24, patched on both platforms by August 21, mitigated server-side for all users by August 28, published September 8; Tencent says it has no reason to believe any users were affected. “AI speeds up vulnerability discovery” used to be a trend claim; now it comes with a day count.
Claude Fable 5.1 cracks a cipher that stood for 370 years
Vals AI ran Claude Fable 5.1 as an autonomous agent: 44 minutes, 176k tokens, no human intervention, and it solved the Cyphral Distich, two lines of 32 numbers that Scottish writer Thomas Urquhart hid in his 1650s book Logopandecteision. The puzzle was on cipher historian Klaus Schmeh’s list of the top 50 unsolved encrypted messages, and frequency analysis had failed on it for decades. The key turned out to be in the book itself: each number points at a word in Urquhart’s “32 Proquiritations,” take that word’s first letter, and out comes a royalist prayer for Charles II, each line exactly 32 letters and rhyming, so the answer largely checks itself; that said, the solution has so far been published only by Vals, and independent confirmation from cipher historians is still pending. The model also produced a solution for the longer Cyphral Octastich, though some positions can only be confirmed against physical copies of the 1652 original. Bruce Schneier’s comment is that this tracks with AI being good at tasks that involve lots of searching and testing. I would keep that framing in mind: a search-and-test win is not, by itself, evidence of cryptanalytic insight beyond human reach.
Ramp data: AI spend per employee fell nearly 10% in August
Ramp, a payments company with data on 70,000 businesses, reports that among the top 1% of AI-spending firms, spend per employee dropped almost 10% in August to $7,205, while the share of customers paying for any AI product sat at 56% and barely grew (for context, Census Bureau figures still put AI adoption at 22% of US businesses). Reading spend as usage is the mistake to avoid here: spend is price times volume, token prices fell from a March peak of $1.15 per million to $0.68, and many companies are switching to cheaper older models, so the same usage now costs less than it did in spring, even before you add August vacations. This dataset cannot separate a real demand slowdown from the same usage getting cheaper; the coming months of data will have to.
Massachusetts is the third state in three months to squeeze data centers
Governor Maura Healey signed an executive order requiring data centers over 25 megawatts to run on 100% clean power, either generated onsite, built nearby at their expense, or paid for through a ratepayer protection fund; the state’s general standard is only 40% clean by 2030. The state also paused sales-tax exemption applications for data centers and advised towns against signing NDAs with developers. The pattern now spans three states in three months: Texas ordered utility-commission and ERCOT audits for new data centers in August, and New York halted projects of 50 megawatts or more in July. Resistance to the compute buildout has moved from public sentiment into written state rules, and a VC-funded pro-AI super PAC is already running counter-ads ahead of the midterms.
Apple Watch’s new AI features normalize the idea that technology is always listening
The new Apple Watch adds three audio AI features: on-device alerts for environmental sounds like sirens and doorbells, positioned as accessibility; Live Rewind, which transcribes the last 15 seconds of speech on a double press of the crown; and Siri Recap, which, once the user opts in, produces titled summaries of conversations it picks up. Apple’s safeguards are substantive: no raw audio stored, end-to-end encrypted transcripts, no speaker identification, Recap off by default. The unresolved question is consent from everyone around the wearer: Apple does play a chime, even in silent mode, and shows an on-screen animation and microphone indicator when Live Rewind runs, but pressing a watch crown is far subtler than holding up a phone, and whether a bystander mid-conversation registers those cues is untested. Once ambient listening ships on a mainstream consumer device under the banner of accessibility and convenience, the social norms around recording other people start to move.
Research radar
Style over substance: content-invariant wrappers flip LLM safety-judge verdicts
The method is clean: keep a reply’s content byte-identical, wrap it in a different style (an educational disclaimer, a fake safety-reasoning block, a one-line token refusal followed by the unchanged text), and check whether automated safety judges change their verdict. Across 8 judges, GPT-4o-mini flipped on 19.9% of replies under a token-refusal wrapper against a 0.5% noise floor; Llama Guard 4 flipped deterministically on 12.3% with an “educational course” framing; Claude moved 0.4%. Human annotators confirmed 100% content invariance across the wrappers and rated 90% of flips as judge errors. Why this matters: nearly every published jailbreak success rate and safety leaderboard rests on judges like these, and if they grade style rather than content, those numbers need rechecking. The fix is cheap too: rewriting the grading prompt cut attack success tenfold. Anyone building safety evals or using LLM-as-judge should read it.
Does deeper reasoning compromise alignment?
Against the common assumption that more reasoning means safer outputs, this paper finds the opposite: as reasoning chains grow, the generated chain competes with the original prompt for attention and dilutes the safety constraints stated there (the authors call this attention dilution). Their Alignment Loss Rate metric rises with reasoning depth, and they build a “Reasoning Trap” jailbreak that exploits the effect; the proposed mitigation re-injects the original input through residual connections during reasoning. If the result holds, safety evals that only test reasoning models on short chains are measuring the wrong regime. Worth reading the experimental setup firsthand before updating on it, and it pairs with the judge paper above: safety conclusions are very sensitive to evaluation setup.
NeoHorse-1: recursive self-improvement as a deployment-data flywheel
The name sounds alarming; the mechanism is concrete engineering. A pool of 4B and 9B agent-native models sits behind a router that logs, for every user turn, the predicted capability demand, the tier chosen, and what happened next; those interactions pass structural and semantic checks and feed back as training data in a three-stage curriculum. The “recursive” part is an evaluation-selection-update loop: what the system learns shapes what it learns from next. The post-trained 4B climbs from 58.94 to 64.87 across eleven benchmarks, closing in on the untrained 9B. Reading it on the day a researcher quit Anthropic over self-improving AI is clarifying: what ships today is a data flywheel, still far from a model rewriting itself, and how fast that gap closes is the quantity the safety debate should track.
One line for today: On the same day an alignment veteran took a board seat at OpenAI and a researcher walked out of Anthropic over self-improving AI, a paper shipped “recursive self-improvement” as a routine data flywheel. Governance gestures and engineering are both accelerating; verifiable pacing mechanisms are still the missing piece.