FT: Anthropic’s strongest model is winning benchmarks and losing budgets

The Financial Times, citing people familiar with the numbers, reports that Anthropic’s annualized revenue reached $65 billion in July, yet Claude Fable 5, its most capable model, accounts for only about 11% of what enterprise customers spend on Anthropic products. Average prices customers actually pay have fallen nearly 25% since mid-July. The mechanism is model routing: companies send routine tasks to cheap models and save the flagship for the hardest problems (Ramp’s CTO describes doing exactly this: lowering default reasoning levels company-wide and barring automated pipelines from frontier models). Anthropic itself is feeding that shift: Claude Opus 5, released July 24, is officially positioned as “close to the frontier intelligence of Claude Fable 5 at half the price,” with an effort dial that lets users trade capability against cost. Revenue is rising while the flagship’s share shrinks, which tells me frontier capability is turning into a premium niche and the volume business lives in the mid-tier. The leaderboard winner and the invoice winner are, increasingly, different models.

TechCrunch surveys where the case law stands, and the key distinction is now visible: courts examine how a company obtained its training data and what it did with the data as two separate questions. In the Anthropic case, Judge William Alsup found the act of training to be fair use, comparing the model to a reader who studies books in order to write something new rather than to replicate them; what cost Anthropic a $1.5 billion settlement with authors was sourcing books from pirate “shadow libraries.” Ross Intelligence lost on the other axis: it used Thomson Reuters legal-database content to build a directly competing legal research product, and the court found no transformative purpose. Thaler v. Perlmutter adds that fully AI-generated works get no copyright at all. So “buy the books, scan them, train on them” is an outline sketched from a handful of cases, not a settled rule: Alsup’s ruling was a single district-court decision, and the settlement means it will never reach an appeals court to become binding precedent. Lawful sourcing and non-substitutive use are two independent ways to lose, and whether the outline holds for every company and every content type is a question no one can answer yet.

Flock’s reforms come with a built-in bypass

Flock Safety operates license plate cameras and drones across thousands of US communities. The Washington Post documented 46 cases of police officers misusing the system, including stalking ex-partners, and the backlash now spans both parties: Democratic candidate Abdul El-Sayed attacked the mass rollout of Flock cameras, while three House Republicans filed a bill to bar federal purchases of such systems. CEO Garrett Langley is calling for a national “compromise” between privacy and safety, and the company points to reforms: default data retention cut from 30 days to 7, and a case code now required before searches. The retention limit, however, can be extended through a feature called Evidence Mode, which preserves searched data for ongoing investigations, and the ACLU asks whether this is reform or public relations. A simple test for reforms like this: measure the width of the exception path. When defaults tighten but the override stays open, the rules only bind people who already follow rules, and the 46 misuse cases came from the people who don’t.

Fabien Sanglard published the agent.md he uses to keep coding agents in line

Fabien Sanglard, the developer known for dissecting game engine source code, published his personal agent.md. An agent.md is a rules file that a coding agent reads at the start of a session, a standing style guide that saves you from repeating the same corrections over and over (the exact filename varies by tool: AGENTS.md, CLAUDE.md, and similar conventions coexist). His rules are specific: minimal comments and replies, function names under 30 characters, fields private by default, a failing test before any bug fix, imperative commit subjects under 50 characters. The list itself is personal taste; the practice is what’s worth copying: turn the review comments you keep making into configuration, and spend your own attention on architecture and design. He is also blunt about the limits: rules or no rules, LLM output still needs line-by-line verification.

One line for today: leaderboards rank peak capability, invoices rank cost per task, and the two rankings crown the same model less and less often.