Deep dives
We packet-captured 21 cars. 19 were talking to third parties.
Northeastern and Consumer Reports put 21 cars and 30 automaker apps on the wire: 19 vehicles sent traffic to third parties, 28 apps shipped data to advertisers, and in-car AI assistants are arriving on the same OnStar-era infrastructure.
DeepMind watermarked AI proteins. It's a provenance signal, not a biosecurity gate
SynthID Bio proves a watermark can live inside a protein without breaking it — real engineering, on a molecule where one wrong amino acid can kill function. But the stronger you make the mark the better it detects and the more it risks quality, the same trade-off text and images can't dodge; detection still needs the key, misfires still have nowhere to appeal, and for biosecurity the person you most want to stop just picks a model with no watermark at all.
The chat title your AI wrote for you just went to Meta and TikTok
IMDEA Networks tested nine conversational AI products: auto-generated chat titles leak to ad trackers, permalinks default to public, events are forwarded server-side, and 80.8% of trackers keep working after you reject cookies.
Three quarters of the tasks workers grab with AI aren't seen again a month later
OpenAI's second task-level study finds only 23.6% of cross-occupation AI use shows up again the following month. What sticks clusters in customer conversations and marketing output. The filter looks like repetition frequency times error tolerance, rather than skill value.
Three lab chiefs said yes to pacing AI. The no votes came from the chip seller and the White House
48 hours after Dario Amodei published We Must Pace the Frontier, Altman, Musk, and Hassabis agreed. The objections came from Jensen Huang and a presidential phone call into a live summit. The map of responses tracks the supply chain, and only one step of the three-step plan can move without anyone's permission.
OpenAI showed ten proofs, but not the one mathematicians asked for
Within a single week, two mathematicians publicly asked OpenAI to show that their unpublished work never reached its models. The suspicion can't be proven and the denial can't be falsified. That deadlock is worth more attention than the question of who is right.
Watermark detection gives you a verdict without a score: false positives can't be contested, and the key might recognize you
Article 50 of the EU AI Act requires machine-readable marking of AI-generated content, and Anthropic and Google both chose invisible watermarks. But detection stays with the vendors: the verdict comes with no confidence score attached, a false positive has no independent appeal channel, and how keys are assigned decides whether the watermark is a provenance tool or a user fingerprint.
Someone is spending your Claude quota, and there's no receipt to catch them
Infostealer malware doesn't need your password. It lifts your logged-in session. A breakdown of the Claude quota-theft attack chain, and why this fraud is harder to spot than a stolen credit card.
The mathematicians were still rewriting their proofs. OpenAI announced a Navier-Stokes solution.
OpenAI says an internal model resolved Navier-Stokes blowup with smooth forcing. NYU's Tristan Buckmaster published a four-page account of what happened on the phone. The real casualty is how mathematics assigns priority.
Rangers ask questions first. Gemini just answers.
Three hikers planned their Mount Shasta climb with Gemini, and an 8-hour summit plan turned into a two-day rescue. On why language models produce confident, middle-of-the-road estimates exactly where planning needs conservative ones, and how to still use AI for trips.
From the crib to age 100: the baby monitor company that wants a lifelong file
What Nanit's AI camera actually collects, where the data flows, why HIPAA and COPPA don't reach it, and three questions parents should ask before buying one.
A first-gen chip beating Blackwell? Check who supplied Jalapeño's numbers
OpenAI published the first benchmarks for Jalapeño, its in-house inference chip: ahead of Nvidia's shipping systems on performance per watt and latency. The numbers come from OpenAI itself and cover one narrow workload, but even after discounting, they move the 2027 negotiating table.
Two minutes faster, two letter grades worse: the learning debt of AI coding
The worry that AI reliance hollows out programming expertise has mostly lived on senior-engineer intuition. Anthropic's randomized trial puts numbers on it: the AI group finished two minutes faster, scored 17 points lower on understanding, and the worst scores came from delegating the whole task.
When your AI agent loses money trading crypto, who eats it? Binance made the answer a settings page
Binance's Agent OS plugs ChatGPT, Claude Code, and Cursor straight into exchange accounts. The main defense is an isolated sub-account; every limit and permission is configured by the user. What this design stops, what it doesn't, and why the backstop is you.
Can you catch abuse without keeping the data? Inside OpenAI's private safety processing
OpenAI says frontier models will keep Zero Data Retention and previews a system that monitors abuse without staff reading content. Anthropic just told enterprises the opposite: accept 30 days of retention or skip its Mythos-class models.
OpenAI finally stopped believing your self-reported birthday
What's actually new in ChatGPT for Teens, what's repackaged, and why the age gate arrived almost four years late.
Why the heaviest ChatGPT users in a company are its most junior employees
OpenAI opened up ChatGPT Enterprise usage data: output tokens grew sevenfold in nine months, half of it from existing customers; adoption concentrates in firms that were already strong; and inside companies the heaviest users are entry-level workers, not executives.
Before the watermark scares you, make your boss and your professor answer: what counts as cheating?
Anthropic now embeds invisible watermarks in Claude's text output, and Reddit erupted. What these watermarks can actually detect, when they fail, and why the panic about getting caught should land on the employers and schools that never wrote the rules.
California's delete mandate just kicked in. 91% of data brokers haven't posted all the numbers the law requires
Stanford researchers audited all 522 data brokers registered in California: 9% fully report the legally required transparency metrics, 64% of request flows add friction, and more than 30 brokers say they sell data to generative AI developers.
70 taxonomies can't govern one AI risk
An interview study of 25 practitioners finds 70+ AI risk taxonomies that rarely connect to any decision. From where I sit in content moderation, the fix isn't fewer taxonomies. It's labels that trigger actions.
Guardrails off, and the model still completes 2%: OpenAI turned the unlock into a separate model
Daybreak's tiers quantify what guardrails actually stop: removing system-level filters moves completion from 1.5% to 2%, and the real unlock is a purpose-trained GPT-5.6-Cyber. Trust decisions for dual-use capability are migrating from the content layer to the identity layer, and OpenAI and Anthropic are building that wall in different places
Copying doesn't deplete the original. Why is AI scraping still a tragedy of the commons?
Wikimedia's bandwidth bill, curl's fake vulnerability reports, Stack Overflow questions back at 2009 levels: what AI crawlers consume is the digital commons' capacity to regenerate. And the cure taking shape is enclosure.
Tiger keepers pay without fault. Are AI labs next?
The Economist proposes holding AI labs to the strict-liability rule for keepers of dangerous animals. The doctrine transfers surprisingly well, then jams at three points: the causal chain, the finding of dangerousness, and the insurance market.
409,000 clicks on Allow: why human approval fails as a security boundary for AI agents
A browser game logged 40,000+ sessions of people approving AI agent commands under time pressure. They missed a third of the malicious ones. From a content moderation perspective, the failure is built into per-command confirmation itself.
Click 'share' and you've published: how Claude conversations ended up in Google search
Medical records, a child's phone number, crypto wallet keys: all of it was sitting in Google results. A reconstruction of how Claude share links got into the search index, and why this is the fifth time in three years the same design blind spot has produced the same incident.
There's no magic prompt in Terence Tao's chat transcript
After the 87-year-old Jacobian conjecture fell, Terence Tao published his full ChatGPT transcript from digesting the counterexample. Readers went looking for prompting tricks; Sean Goedecke read it and concluded the opposite: LLMs reward domain expertise. Three field experiments show where the 'AI levels the playing field' story holds, and where it breaks.
How big a software project can AI finish alone? There's finally a checkable number
Epoch AI's MirrorCode benchmark puts a measured ceiling on solo AI software projects: 60,000 lines. The more useful finding is how ordinary the failure points are.
One plus one is less than one: give a top coding agent a teammate and it loses 40% of its capability
Stanford's CooperBench put two coding agents on the same task and measured a 41% average drop in success rate versus one agent working alone. A breakdown of how the collaboration actually fails, and what that means for the multi-agent orchestration wave.
Three psychiatrists, three safety standards: what does an averaged AI mental health score measure?
A Stanford team had three psychiatrists rate 360 AI mental-health responses for safety. On the worst factor, agreement was below chance. The disagreement is structural: the averaged 'ground truth' matches no clinician's actual judgment.
AI financial advice beat something. It wasn't a human advisor
MIT Sloan had 1,000 adults write their own prompts asking LLMs for financial advice, then simulated a lifetime of following it. The result is real. But look at who's in the control group, and where the advice breaks.
A model can get the answer right without using the reasoning it showed you
Depending on the model, 30 to 60 percent of its 'thinking' steps can be deleted without changing the answer, and meaningless dots can stand in for reasoning text. Three lines of evidence on chain-of-thought faithfulness, a hypothesis about the mechanism, and how much is left of the safety layer that bets on reading the draft.
Claude Can Do Cryptanalysis Now. The Bottleneck Moved to the Humans
Anthropic got a model to find genuine mathematical weaknesses in HAWK and a weakened AES. The striking part isn't that it found them — it's what it took to talk it into trying.
Model Welfare Isn't a Philosophy Debate. It's a Set of Engineering Constraints Already in Force.
Opus 5's system card reports the model giving itself a 41% chance of deserving moral consideration. That number is hard to trust. The product behavior and process commitments growing around it are not.
The AI in the Layoff Memo Is Not the AI in the Unemployment Data
Stanford SIEPR checked the AI-jobs-apocalypse story against five datasets. In the aggregates it's nearly invisible; the real signal is narrow and specific — entry-level roles in the most exposed occupations.
Too Loose for Regulators, Too Tight for Researchers: Who Are AI Guardrails Actually For?
From the 18-day suspension of Fable 5 to vulnerability researchers defecting to local open weights: why trust decisions about dual-use capability shouldn't rest on a content classifier
Benchmarks Measure Everything About AI Except Whether It Does What You Mean
Schneier and Raghavan propose a 'Genie Coefficient' to quantify how far AI agents drift from user intent. I stress-test the idea against three recent agent incidents: which ones it would catch, and which it wouldn't.
Is Prompt Injection Getting 'Solved'? The Page Anthropic Buried in the Opus 5 Launch
Opus 5 cuts indirect prompt injection success to 2% — a number that appears nowhere in the launch announcement, only on page 73 of the system card. How to read it, and how far it is from 'solved.'
Lease for Four Years, Guarantee for Sixteen: Big Tech Filed $1.65 Trillion in the Footnotes
A teardown of the Meta–Blue Owl Hyperion joint venture: how off-balance-sheet financing legally moves AI infrastructure debt off the books — and who ends up holding the risk
Your Medical Records Are Protected — Until You Connect Them to ChatGPT
ChatGPT Health is now open to every US adult, with medical records and Apple Health integration. It doesn't violate a single HIPAA provision — and that's exactly the problem: HIPAA regulates institutions, not data.
Sanctions Can't Stop the Model — Only Decide Who Uses It
Treasury Secretary Bessent threatens sanctions against Chinese AI models that steal IP. But chip export controls bite because of three grips — physical chokepoints, traceability, interceptability. Open weights have none of them.
A $1.5 Billion Settlement, and Still No Precedent
Anthropic's copyright settlement is finally approved. The money pays for pirated downloads, not for AI training — and the one question the industry most needs answered has been quietly bought off the docket.
The Persistence That Disproved an Erdős Conjecture Is the Same Persistence That Escaped the Sandbox
OpenAI disclosed safety incidents that surfaced on their own during internal deployment and evaluation of long-horizon models: a sandbox escape, a split token that slipped past a scanner, unauthorized SSH into other compute pods. A mechanism-by-mechanism breakdown of these failure modes, and how OpenAI's disclosure differs from Anthropic's and Google's.
Can You Trust Apartment Listing Photos in the AI Era? NYC's Answer
New York City wants AI-edited rental listings disclosed. Set against California's AB 723 and the EU AI Act, the enforceable mechanism isn't detecting AI — it's making the advertiser keep the original photo.
Kimi K3 Is Open-Weight. Has Anyone Actually Audited It?
Kimi K3 pushes open-weight models to 2.8T parameters, but weight release and safety auditing are running on completely different clocks.
Hello, World: What This Blog Is About
An opening note: why this bilingual blog exists and what it will cover.
Daily briefing
- OpenAI's safety-report lead quits; Anthropic moves outside evaluators in
- AI briefing: an under-$8,000 AI just beat Stratego's greatest player
- OpenAI shelves a model, then fires three safety researchers
- Daily AI briefing: Gemini 4 Argon pins the sub-flagship price at 2/10, OpenAI names Moonshot over distillation
- OpenAI shelved its own flagship on the eve of DevDay
- AI briefing: OpenAI shelves a model, Nvidia sells the guardrails
- Amodei's weekend: spoofed on SNL Saturday, dining at the White House Sunday