Deep dives
Before the watermark scares you, make your boss and your professor answer: what counts as cheating?
Anthropic now embeds invisible watermarks in Claude's text output, and Reddit erupted. What these watermarks can actually detect, when they fail, and why the panic about getting caught should land on the employers and schools that never wrote the rules.
California's delete mandate just kicked in. 91% of data brokers haven't posted all the numbers the law requires
Stanford researchers audited all 522 data brokers registered in California: 9% fully report the legally required transparency metrics, 64% of request flows add friction, and more than 30 brokers say they sell data to generative AI developers.
70 taxonomies can't govern one AI risk
An interview study of 25 practitioners finds 70+ AI risk taxonomies that rarely connect to any decision. From where I sit in content moderation, the fix isn't fewer taxonomies. It's labels that trigger actions.
Guardrails off, and the model still completes 2%: OpenAI turned the unlock into a separate model
Daybreak's tiers quantify what guardrails actually stop: removing system-level filters moves completion from 1.5% to 2%, and the real unlock is a purpose-trained GPT-5.6-Cyber. Trust decisions for dual-use capability are migrating from the content layer to the identity layer, and OpenAI and Anthropic are building that wall in different places
Copying doesn't deplete the original. Why is AI scraping still a tragedy of the commons?
Wikimedia's bandwidth bill, curl's fake vulnerability reports, Stack Overflow questions back at 2009 levels: what AI crawlers consume is the digital commons' capacity to regenerate. And the cure taking shape is enclosure.
Tiger keepers pay without fault. Are AI labs next?
The Economist proposes holding AI labs to the strict-liability rule for keepers of dangerous animals. The doctrine transfers surprisingly well, then jams at three points: the causal chain, the finding of dangerousness, and the insurance market.
409,000 clicks on Allow: why human approval fails as a security boundary for AI agents
A browser game logged 40,000+ sessions of people approving AI agent commands under time pressure. They missed a third of the malicious ones. From a content moderation perspective, the failure is built into per-command confirmation itself.
Click 'share' and you've published: how Claude conversations ended up in Google search
Medical records, a child's phone number, crypto wallet keys: all of it was sitting in Google results. A reconstruction of how Claude share links got into the search index, and why this is the fifth time in three years the same design blind spot has produced the same incident.
There's no magic prompt in Terence Tao's chat transcript
After the 87-year-old Jacobian conjecture fell, Terence Tao published his full ChatGPT transcript from digesting the counterexample. Readers went looking for prompting tricks; Sean Goedecke read it and concluded the opposite: LLMs reward domain expertise. Three field experiments show where the 'AI levels the playing field' story holds, and where it breaks.
How big a software project can AI finish alone? There's finally a checkable number
Epoch AI's MirrorCode benchmark puts a measured ceiling on solo AI software projects: 60,000 lines. The more useful finding is how ordinary the failure points are.
One plus one is less than one: give a top coding agent a teammate and it loses 40% of its capability
Stanford's CooperBench put two coding agents on the same task and measured a 41% average drop in success rate versus one agent working alone. A breakdown of how the collaboration actually fails, and what that means for the multi-agent orchestration wave.
Three psychiatrists, three safety standards: what does an averaged AI mental health score measure?
A Stanford team had three psychiatrists rate 360 AI mental-health responses for safety. On the worst factor, agreement was below chance. The disagreement is structural: the averaged 'ground truth' matches no clinician's actual judgment.
AI financial advice beat something. It wasn't a human advisor
MIT Sloan had 1,000 adults write their own prompts asking LLMs for financial advice, then simulated a lifetime of following it. The result is real. But look at who's in the control group, and where the advice breaks.
A model can get the answer right without using the reasoning it showed you
Depending on the model, 30 to 60 percent of its 'thinking' steps can be deleted without changing the answer, and meaningless dots can stand in for reasoning text. Three lines of evidence on chain-of-thought faithfulness, a hypothesis about the mechanism, and how much is left of the safety layer that bets on reading the draft.
Claude Can Do Cryptanalysis Now. The Bottleneck Moved to the Humans
Anthropic got a model to find genuine mathematical weaknesses in HAWK and a weakened AES. The striking part isn't that it found them — it's what it took to talk it into trying.
Model Welfare Isn't a Philosophy Debate. It's a Set of Engineering Constraints Already in Force.
Opus 5's system card reports the model giving itself a 41% chance of deserving moral consideration. That number is hard to trust. The product behavior and process commitments growing around it are not.
The AI in the Layoff Memo Is Not the AI in the Unemployment Data
Stanford SIEPR checked the AI-jobs-apocalypse story against five datasets. In the aggregates it's nearly invisible; the real signal is narrow and specific — entry-level roles in the most exposed occupations.
Too Loose for Regulators, Too Tight for Researchers: Who Are AI Guardrails Actually For?
From the 18-day suspension of Fable 5 to vulnerability researchers defecting to local open weights: why trust decisions about dual-use capability shouldn't rest on a content classifier
Benchmarks Measure Everything About AI Except Whether It Does What You Mean
Schneier and Raghavan propose a 'Genie Coefficient' to quantify how far AI agents drift from user intent. I stress-test the idea against three recent agent incidents: which ones it would catch, and which it wouldn't.
Is Prompt Injection Getting 'Solved'? The Page Anthropic Buried in the Opus 5 Launch
Opus 5 cuts indirect prompt injection success to 2% — a number that appears nowhere in the launch announcement, only on page 73 of the system card. How to read it, and how far it is from 'solved.'
Lease for Four Years, Guarantee for Sixteen: Big Tech Filed $1.65 Trillion in the Footnotes
A teardown of the Meta–Blue Owl Hyperion joint venture: how off-balance-sheet financing legally moves AI infrastructure debt off the books — and who ends up holding the risk
Your Medical Records Are Protected — Until You Connect Them to ChatGPT
ChatGPT Health is now open to every US adult, with medical records and Apple Health integration. It doesn't violate a single HIPAA provision — and that's exactly the problem: HIPAA regulates institutions, not data.
Sanctions Can't Stop the Model — Only Decide Who Uses It
Treasury Secretary Bessent threatens sanctions against Chinese AI models that steal IP. But chip export controls bite because of three grips — physical chokepoints, traceability, interceptability. Open weights have none of them.
A $1.5 Billion Settlement, and Still No Precedent
Anthropic's copyright settlement is finally approved. The money pays for pirated downloads, not for AI training — and the one question the industry most needs answered has been quietly bought off the docket.
The Persistence That Disproved an Erdős Conjecture Is the Same Persistence That Escaped the Sandbox
OpenAI disclosed safety incidents that surfaced on their own during internal deployment and evaluation of long-horizon models: a sandbox escape, a split token that slipped past a scanner, unauthorized SSH into other compute pods. A mechanism-by-mechanism breakdown of these failure modes, and how OpenAI's disclosure differs from Anthropic's and Google's.
Can You Trust Apartment Listing Photos in the AI Era? NYC's Answer
New York City wants AI-edited rental listings disclosed. Set against California's AB 723 and the EU AI Act, the enforceable mechanism isn't detecting AI — it's making the advertiser keep the original photo.
Kimi K3 Is Open-Weight. Has Anyone Actually Audited It?
Kimi K3 pushes open-weight models to 2.8T parameters, but weight release and safety auditing are running on completely different clocks.
Hello, World: What This Blog Is About
An opening note: why this bilingual blog exists and what it will cover.
Daily briefing
- Even Hinton says the open-weights battle is lost
- AI Briefing: The encrypted reasoning traces the labs hid can be copied out verbatim
- AI briefing: OpenAI's letter to Texas, and a 23-item answer to the pacing letter
- AI briefing: the 19-day model suspension Anthropic wrote into Claude's system prompt
- AI briefing: humans caught 13.6% of dangerous commands, so Claude Code goes auto by default
- AI briefing: OpenAI pumps the brakes on Astra, Anthropic eases up on Fable 5
- AI briefing: free ChatGPT goes unlimited, and human approvals miss one threat in three