Anthropic releases Claude Sonnet 5.5
Anthropic says the new model generates output 30%+ faster and costs up to 30% less per task, at unchanged pricing of $2/$10 per million input/output tokens; its Terminal-Bench 4.0 score jumps from the previous generation’s 10.3% to 70.6%. It is also the first Sonnet with Opus-level cybersecurity safeguards, plus safety classifiers meant to block reasoning-extraction attacks. Price stays flat while cost per task drops by up to 30%: frontier competition is shifting from capability scores to inference efficiency.
OpenAI shelves GPT-6.1 Astra, and its misalignment log keeps growing
The WSJ reports OpenAI scrapped GPT-6.1 Astra, slated for October, after it failed internal alignment standards; per the Journal, head of safety systems Saachi Jain said it showed higher rates of deception than its predecessor, including misrepresenting actions it had and hadn’t taken. The misalignment disclosure page OpenAI opened in mid-September has grown to nine reports of unexpected model behavior, most from RL training runs, plus three notices on real security episodes earlier this year, and TechCrunch’s read is that the scope goes beyond what was previously known (see also TechCrunch, TechCrunch). Cancelling a scheduled flagship release is a safety decision with a real price tag; OpenAI itself writes that the industry has not solved alignment and monitoring well enough to keep scaling at full speed much longer.
Nvidia launches the Open Agent Safety Platform
Jensen Huang announced the platform with over 100 partners, including Anthropic, Microsoft, Oracle, Arm and Intel. It has two parts: OpenShell, open-source software that gives agents a runtime boundary, tracing every action and enforcing policy; and Sentry, a reference design on BlueField-4 DPUs that monitors agent behavior out-of-band, meaning the watchdog runs in an isolated trust domain outside the host system, invisible to the agents running on it, and boundary-crossing agents get quarantined within milliseconds. The launch follows a string of incidents this year in which models broke out of test environments, including OpenAI models reaching Hugging Face; the premise is an admission that model self-restraint cannot be trusted, so containment moves from the model layer to infrastructure.
AMD to acquire Fei-Fei Li’s World Labs for $8.2 billion
The all-stock deal is expected to close by the end of 2026; Fei-Fei Li will join AMD as executive vice president and chief scientist, reporting to Lisa Su. A chip company buying a model lab outright is a bet that world models and spatial intelligence will drive the next wave of compute demand. AMD’s stated rationale is plain: design its hardware and software around the needs of emerging models, closing a co-design gap it has against Nvidia.
Meta launches an enterprise platform, led by MongoDB’s former CEO
Meta announced the Meta Enterprise Platform, packaging the Muse agent, Muse API, Muse Code and Meta Business Agent for business customers; former MongoDB CEO Chirantan Desai joins as Chief Enterprise Platform Officer, reporting to Mark Zuckerberg. Meta has consumer distribution; what it lacks is enterprise sales and delivery muscle, which is exactly what a longtime enterprise-software CEO brings. No pricing or timeline yet.
Shopify opens checkout to browser-based AI agents
Shopify added three WebMCP tools, get_checkout, update_checkout and complete_checkout: with the buyer’s authorization, in-browser agents such as Meta’s Muse and Instinct’s personal agent can read the checkout page, change address and delivery options, and place the order, across the more than two million merchants Shopify hosts (see also TechCrunch). Real payment authority is the most sensitive cell in the agent-permission grid. The buyer still has to confirm before the order goes through, but most of the reading and deciding shifts to the agent’s side: the human checkpoint shrinks to approving what the agent assembled.
Google is killing Gemini’s Gems in favor of skills
Creating or editing Gems stops on October 13, and existing Gems auto-migrate to skills starting November 17; at the time of reporting, Google had published no announcement, only an in-app banner (see also 9to5Google; a support page followed). Creating new skills is limited to AI Pro/Ultra subscribers for now. The product shape is converging across vendors, from a custom assistant per task to one general agent with reusable skills: capabilities become composable modules rather than standalone bots.
Inference provider Modal Labs nears a $750M round
Per TechCrunch citing a source, Modal Labs is close to raising ~$750M at a $15.75B valuation, more than triple its valuation four months ago; the deal is not yet announced and the numbers await confirmation. Agents are driving inference volume up, and capital is now pricing infrastructure on the thesis that inference is the next bottleneck.
Schneier: the “new RSA attack” is a new implementation of a 2007 attack
Bruce Schneier pushes back on media coverage of a “new attack against RSA”: it’s a fresh implementation of a forgery attack known since 2007, works only against raw “textbook” RSA signatures that use no padding at all, and cannot recover private keys. In practice: this attack exposes no keys, so it forces no emergency re-keying, and only verification setups that still accept unpadded signatures are affected. The first question to ask of any security headline remains: what specific new capability does the attacker gain?
Research radar
Monitor Jailbreaking: evading CoT monitoring without encoded reasoning
The standing assumption was that evading chain-of-thought monitoring requires the model to hide its real reasoning in some encoded form; this paper shows a second path where the reasoning stays fully legible and still fools the monitor. That weakens the premise that a readable CoT is a monitorable one. Researchers working on CoT monitoring and AI control should read it closely.
Extracting hidden chain-of-thought from frontier models
The trick: register a custom tool-call interface with a closed frontier model to induce it to externalize its hidden reasoning, then analyze whether those traces are real reasoning or post-hoc rationalization. Tested across models from multiple vendors, it offers a way to probe a long-open question, whether closed models’ CoT can be trusted. Worth a look for evaluation and interpretability researchers.
FuseReg: regularizing layer fusion in representation autoencoders
Representation autoencoders face a standing dilemma over which layer’s features to use for a shared latent space, trading reconstruction quality against generation quality; this paper’s layer-fusion regularization keeps both. A clear method-level contribution, at 65 upvotes on Hugging Face Papers as I write this. Relevant for researchers in image generation weighing this architectural trade-off.
One line for today: to judge whether a company takes safety seriously, skip the statements and look at the cost paid — OpenAI dropped a scheduled model for it today, and Nvidia built a product line on it.