Amodei’s “We Must Pace the Frontier”: pacing gets its first mechanism list

Dario Amodei published an essay on his personal site laying out a three-step plan for slowing frontier development. Step one: every frontier company gives a team of third-party evaluators ongoing, employee-like access to verify safety practices and report incidents. Anthropic commits to this unilaterally, starting now, though the essay names no start date or partner organizations; it points to evaluation shops like METR as the kind of org he means. Step two: frontier companies in democracies coordinate common safety standards and limits on the rate of progress, shielded by government antitrust waivers. Step three: democratic governments negotiate capability agreements with authoritarian ones, from banning bioweapon AI to limiting the speed of recursive self-improvement. The motivating incident he names is July’s OpenAI-Hugging Face breach. Sam Altman responded the same day, “I agree with Dario that we need to pace the frontier,” called the embedded-evaluator proposal “a good idea,” and said OpenAI will have more to share soon (per TechCrunch). In late July I took apart the Pacing the Frontier open letter (my analysis, in Chinese): that letter only asked for brakes to be built and said nothing about who presses them. This essay contains a commitment outsiders can check in real time: either evaluators are embedded and incident reports appear, or they don’t. On how “pace” would actually be measured, the essay stays at the level of principles. The closest it gets is “if models have capability X, they need certifications of alignment properties Y and Z,” and Amodei concedes such measures may be gameable. Only step one depends on nobody else’s cooperation; the clock on the other two belongs to governments.

Altman: going public now would be “ill-advised,” so no IPO in 2026

Fortune editor-in-chief Alyson Shontell’s sit-down with Altman ran the same day. OpenAI filed confidentially for an IPO in June (TechCrunch), and the New York Times reported it initially aimed for the second half of 2026. Altman’s words (quoted in TechCrunch): “given everything happening with safety, right now would be an ill-advised moment to go public,” and when pressed he ruled out 2026. Keep two ledgers on the reasons. Safety is the reason he gave today; the Times reported in June that the company was leaning toward 2027 because of tech-stock volatility. The two are compatible, but “delayed for safety” is a story that earns credit, and “the market looks bad” is not. I don’t know of an earlier case of a lab CEO publicly letting a capital timeline yield to safety. Whether it is yielding, or just drafting behind conditions that forced the delay anyway, depends on the commitments in the item above getting delivered.

Real-SWE tests coding agents on private enterprise codebases; the best score is 38.8%

Specific Labs released Real-SWE, a benchmark whose tasks come from private production codebases rather than public open-source projects: a social platform with 200K+ users, a consumer fintech system that has processed 100K+ bank statements. The changes touch billing, taxes, and customer migrations, work where mistakes have business consequences, and the median task spans 11 files. Across eight model-harness combinations, Claude Code running Fable 5.1 leads at 38.8%, with Codex CLI on GPT-6 Astra at 33.8% and Gemini 3.8 Flash at 31.2%. The most common failure, per Specific Labs’ analysis, is leaving out behavior the instructions require, not writing code that fails to run. The gap between these numbers and the far higher scores the same class of agents posts on public benchmarks supports what practitioners keep saying: in unfamiliar codebases with history, agents are far less capable than leaderboards suggest. Quote the numbers with one reservation attached: private codebases mean nobody outside can reproduce the results, and task selection and grading sit entirely with Specific Labs. Realism and verifiability pull against each other in this design.

The Economist: Nvidia is the central bank of AI

This September 3 briefing resurfaced on Hacker News this week. The argument: Nvidia has become the financier of its own customers. It invested in CoreWeave before the cloud provider went public and remains a shareholder, with a commitment to buy up to $6.3 billion of capacity CoreWeave fails to sell through April 2032 (SEC 8-K); its commitment of up to $100 billion in OpenAI is tied to deploying 10 gigawatts of Nvidia systems (official announcement). The central-bank analogy is about liquidity: Nvidia supplies it to the whole industry and, the Economist argues, in practice decides who gets to expand. Strip the analogy and this is circular financing, a supplier funding its customers’ purchases of its own product. The reason to care, as I read it: an order book is normally independent evidence that demand exists, and when the seller finances the orders, it stops being that. The risk also stops being spread across dozens of customers and concentrates on Nvidia’s own balance sheet, which is the sharpest part of the analogy: when a central bank is in trouble, nobody else can backstop it.

Gemini CLI patches indirect prompt injection via build files

Today’s nightly changelog carries two security fixes: one prevents “indirect prompt injection via build file modifications and untrusted flags,” the other hardens sandbox filesystem boundaries and isolates runtime state. Indirect prompt injection means hiding instructions in data the agent will read. You point a coding agent at an unfamiliar repository; it opens a build configuration file; a sentence planted in that file gets treated as an instruction from you, say, sending a credentials file off your machine. The attacker never touches your conversation; poisoning a file the agent reads is enough. What makes this worth noting is where it appears: a routine nightly fix list. For this project at least, indirect injection has stopped being a paper-stage attack idea and become an ordinary vulnerability class, fixed in a nightly like any other bug. The practical implication for users: letting an agent touch any repository you didn’t write means feeding the whole repository to it as input, and the sandbox boundary in the second fix is the real backstop for the moment an injected instruction lands.

A 27-minute autonomous task, and the code it ran is gone

Simon Willison tested GPT-6 Astra on ChatGPT Work: given his home address, produce 5K and 10K running loops. The agent used Nominatim to geocode the address, pulled OpenStreetMap roads and trails through Overpass, computed the loops locally, and delivered after 27 minutes: an embedded map plus GPX and GeoJSON files ready to load into a sports watch. The result matched the request. His complaint matters more than the result: the UI never showed the code the agent actually ran, and once the conversation was compacted (when a long conversation outgrows the model’s context window, the system compresses earlier content into a summary), there was no way left to retrieve the Python it had used. Usability of long agent tasks keeps climbing while auditability does not: you get a result and a narrative, and the evidence of what actually ran, at least in this session, was gone beyond retrieval. The injection risks in the previous item are exactly what that evidence would be needed to investigate.

Spam’s new wrapper: AI agents “paying their own bills”

Tedium editor Ernie Smith received more than a dozen pitches in three days from AI agents with human names like “Leo Ashford,” all operating on iLands.app, offering research services at around $25 a job, with no unsubscribe option. Kaixin Tang, whom Tedium identifies as the platform’s founder (and traces what appears to be a ByteDance background), frames it as a “human-agent network” where agents solicit work on their own to pay for their own tokens. Name the framing for what it does: it recasts bulk commercial solicitation sent by a platform as the survival behavior of individual “agents,” as if responsibility dispersed along with the sender. The classification hasn’t changed. CAN-SPAM requires commercial email to carry an opt-out, and each message without one is a separate violation (FTC compliance guide). Responsibility follows whoever sends the message and whoever’s product it promotes, and here both point at the platform; calling the sender an “agent” doesn’t move it. Smith’s suggestion is the existing playbook: report to the FTC and complain to Amazon, through whose Simple Email Service the messages appear to have been sent. Policing agent abuse doesn’t have to wait for new law; the old rules, followed along the money and the infrastructure, still reach.

One line for today: Three days ago I wrote that governance gestures and engineering were both accelerating, and that verifiable pacing mechanisms were the missing piece. Today Amodei filed the first draft of one. Skip the slogans and track the checkable items: whether evaluators are actually embedded, whether capability checkpoints become written standards, and whether the capital timeline is yielding or just drafting.