OpenAI’s safety-report lead resigns: “the time for trial and error is over”

David Robinson, who spent three and a half years at OpenAI, oversaw the system cards (the safety reports shipped with model launches) for 12 frontier releases, and helped draft the Preparedness Framework the company uses to decide when a model is too dangerous to ship without extra safeguards, resigned with an essay in The Atlantic titled “I Quit OpenAI Because Its Culture Is Broken.” His argument: iterative deployment guarantees periodic failures, and frontier labs should run like nuclear plants or busy airports, with layered redundancy and slow, deliberate planning. It follows reports of fired safety researchers and a shelved model release, and it is hard to dismiss: the person calling the process broken is the one who oversaw its safety reporting and helped write its rules.

Anthropic and Accenture each commit $1B to “embedded evaluation”

Anthropic will host external evaluators from Accenture inside the company with employee-level access: they can watch models take shape during training, trace build and deployment decisions, and talk directly to staff, covering model evals, red-teaming, alignment assessments, and safeguard testing, with each company planning to put in at least $1 billion over the next five years. The money is the awkward part: no funding standard exists yet, so Anthropic pays Accenture directly at first, meaning the evaluated party pays its evaluator; Anthropic says it is discussing pilots with nonprofits like METR that would run on the nonprofits’ own funding. Even so, this goes a level deeper than running benchmarks over an API: the object of evaluation becomes the process that builds the model, not the artifact that comes out.

Anthropic’s official guide to getting the most out of Opus 5.5

Anthropic’s usage guide for Opus 5.5 in Claude and Claude Code (by Addy Osmani) picked up a sizable Hacker News thread this week. The advice: give the whole task in one message with an explicit definition of done, drop “think carefully” filler since the model reasons before every reply, use CLAUDE.md to set when it should pause for you versus keep going, have it keep task lists in files on long runs, and parallelize large audits with subagents. Almost none of it is about wording; what you tune is the stopping condition, not the sentence.

AWS answers the data center backlash: no more NDAs, $1B for host communities

AWS CEO Matt Garman responded to community opposition over water use, electricity prices, generator pollution, and the nondisclosure agreements that kept local governments from telling residents about projects before permits were locked in: the company no longer signs NDAs with the government agencies it works with, and a five-year, $1 billion “Built Together” program will fund free community-college degrees plus energy-efficiency upgrades for over 300 schools and 30,000 homes. He also called the water and power complaints “myths” and warned that the 100-plus moratoriums under consideration across the US would cost the country the AI race. Two caveats: the NDA change covers government agencies only and the announcement does not say whether existing agreements end (when Microsoft announced the same policy about half a year earlier, it terminated existing NDAs too), so transparency is turning into the price of admission for building compute rather than a differentiator. (See also TechCrunch.)

Stanford HAI: open weights are not open source

Stanford HAI director James Landay separates two things that get conflated: releasing weights is open distribution, where you can download the file but cannot see the training data, the code, or the decisions behind the model. True open source, in his framing, means the Linux Foundation Model Openness Framework’s “Open Science” tier, which releases training code, training data (or an auditable account of it), and tooling, and he wants universities to lead fully published frontier model development. The stake for research is reproducibility: without data and recipe, “why does the model do this” stays unanswerable, and that is exactly the question alignment work needs answered.

Research radar

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Across 13 pretrained models, tracking reference chains in context (line N cites line M, follow it back step by step) tops out at a reliable depth of 1.4 to 3.6 lines; training a single rank-8 LoRA at one early layer, with everything else frozen, lifts Qwen3-8B from 15.5% to 99% exact accuracy on 24-line chains, and longer training reaches 50 lines. The mechanism is a relay: the LoRA gets each line to pass its chain identity through the middle layers, and frozen attention heads read progressively further up the chain, so the frozen components can carry the computation once a tiny adapter activates it. Worth reading for long-context, interpretability, and post-training researchers: your task may have the same kind of idle depth that one small LoRA unlocks.

Today in one line: The person who oversaw OpenAI’s safety reports quit saying the internal process can’t be trusted; two weeks earlier, Anthropic handed outside evaluators a badge to the building. One deal doesn’t make an industry shift, but the direction of travel is from self-report toward on-site inspection.