Anthropic releases Claude Opus 5.5 with a 20% price cut

Anthropic says Opus 5.5 performs at the level of Fable 5.1 while costing 40% less to run than Opus 5; API pricing drops to $4/$20 per million tokens (20% off) and cache reads fall 60%. The benchmark numbers (66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.0) are Anthropic’s own. Safety gets top billing: METR and Frontier Design ran external alignment testing, and cybersecurity and biology tasks stay restricted behind verification programs — the life-sciences one already in place, the cyber one expanding in the coming weeks. This is Anthropic’s first model since Amodei’s essay calling to pace the frontier ten days ago (my analysis), and the first move after that call is a cheaper, faster rollout.

OpenAI ships GPT-6 Sol and Luna at half the price, 90 minutes later

The launch came about 90 minutes after Anthropic’s, per TechCrunch. VentureBeat has the prices: Sol, aimed at coding and professional work, costs $2/$10 per million tokens, half of GPT-5.6 Sol; Luna, for high-volume tasks, costs $0.10/$0.50, down from $0.20/$1.20. An OpenAI spokesperson confirmed these are permanent prices, not a promotion. Note the exact positioning: Sol’s new rate is precisely half of Opus 5.5’s freshly cut one, so Anthropic’s 20% cut held for an hour and a half before being undercut by 50%.

Pentagon review: AI overreliance contributed to the strike on an Iranian school

Bloomberg, citing officials involved in an unreleased internal Pentagon review, reports that on the war’s opening day (February 28) two Tomahawk missiles hit an elementary school in Minab, Iran, killing more than 150 people, at least 123 of them children; investigators found CENTCOM personnel leaned too hard on Palantir’s Maven Smart System targeting platform, on top of outdated satellite imagery. The failure chain is specific: an analyst had logged signs the building was a school back in 2019, in a system not connected to the main targeting database, and the civilian-harm-mitigation staff who might have caught it had been cut by about 90%, down to a single person at CENTCOM, with no review of the site before the strike. This is a real strike, not a red-team exercise — Amnesty International independently documented it from satellite imagery, video, and witness testimony — and the pattern matches what I wrote in human approval is not a security boundary: the loop failed long before the model did, because the people and channels that could correct the error had already been removed.

GPT-6 Astra breaks an Enigma message unsolved for 21 years

Crypto Cellar Research’s first-hand write-up: deployed by researcher Carter Leffer, GPT-6 Astra picked the July 10, 1941 German Army message MVUEH out of a pool of unbroken ciphertexts (public since 2005), wrote its own Enigma simulator and Bombe software in Python and C++, used the repeated place name “ROSENOW ROSENOW” as a crib, and recovered the key; veteran codebreaker Frode Weierud verified the break. The solve also explains why the message resisted so long: transcription errors in the original, plus a rare turnover of the left wheel at letter 72. As with Claude’s cryptanalysis work I covered in July, what matters is the full autonomous chain of picking the target, building the tools, and verifying the result; historical ciphers are a harmless proving ground, and security researchers should assume the same chain will be pointed at weak modern systems.

Meta admits Muse’s likeness to OpenClaw is no coincidence

Nat Friedman, product lead at Meta Superintelligence Labs, acknowledged on X that Muse was “heavily inspired as a product by OpenClaw”: the workspace file naming matches, and the SOUL.md file that defines an agent’s personality and boundaries is nearly identical in content, though Meta says the code was built from scratch. His explanation: “we thought that Peter got those things exactly right,” referring to OpenClaw creator Peter Steinberger. With Muse sitting at #1 on the US App Store (yesterday’s briefing covered Amazon blocking it from shopping on Amazon.com), the lesson is blunt: product-layer design in agents copies at essentially zero cost, and an independent developer’s good decisions transfer to a trillion-dollar company overnight.

OpenAI publishes priorities and principles for third-party assessments

OpenAI names four areas it wants external assessors to examine: end-to-end safety cases, critical safeguards (jailbreak resistance, misalignment monitors, cyber defenses), capability evaluations for its Preparedness risk categories, and independent investigation of critical misalignment incidents; its principles include pre-registered claims, proportionate access, conflict-of-interest disclosure, and remediation time before publication. It sits on the same thread as Opus 5.5 naming METR as an external evaluator the same day: both labs now build third-party review into their launch story. Two things will decide whether this is real independence. On access, the document does commit to depth — enough for assessors to challenge OpenAI’s assumptions and reach their own conclusions — but qualifies it with legal, security, and IP limits, and allows substitutes where direct access is “impractical.” On the remediation window, it sets no upper bound, so whether it becomes a buffer against critical findings is left open.

Research radar

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

With the model frozen, harness self-improvement lets an agent edit its own prompts, tools, control flow, and memory, and the known failure mode is overfitting: big gains on training tasks that vanish out of distribution. RRSI constrains the loop: the proposer gets an edit budget per candidate and must explore novel trajectories, while a critic-plus-pruner selector drops changes that are marginal, costly, or obsolete, favoring reusable mechanisms over benchmark-specific hacks. Results: up to 14.1 points in-distribution, up to 4.7 points across five out-of-distribution benchmarks, and 30% fewer policy tokens; worth reading if you build self-modifying harnesses and worry about them drifting.

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Video world models tend to forget what they showed you once the camera moves or time passes. WorldCrafter compresses past observations into view-specific memory tokens, retrieved by the requested camera pose and injected before denoising, with no explicit depth correspondences, and it supports minute-scale scene exploration. Reported gains cover long-horizon consistency and camera-control accuracy; relevant if you work on world models or controllable video generation.

One line for today: Ten days after a call to pace the frontier, the answer arrived as a same-day price war; the one checkable indicator of whether the safety commitments still bind is how much access third-party assessors actually get.