The encrypted chain-of-thought from the big three can be lifted verbatim

Researchers from the ELLIS Institute Tübingen and the Max Planck Institute (Panfilov, Shumailov, Geiping, Andriushchenko, and others) published a paper exposing a structural flaw in how Anthropic, OpenAI, and Google hide their models’ chain-of-thought. To keep their reasoning from being copied, these providers don’t store the trace server-side — they encrypt it into a blob, hand it back to the client, and have the client pass it back on the next request. The catch: within a single provider’s ecosystem, those encrypted blobs are interchangeable across sessions, across users, and even across models. The team turned that into a scalable “decryption jailbreak” — feed a strong model’s encrypted reasoning blob to a weaker, less-guarded model from the same provider, and the weaker one decodes the ciphertext into plaintext word for word.

They pulled 6,708 public agent trajectories off GitHub and Hugging Face (each carrying encrypted reasoning blobs), ran their decoding pipeline, and reconstructed 315,320 reasoning blocks — and from real user sessions recovered 704 privacy artifacts, including 62 API keys, 33 passwords, 24 access tokens, and 30 email addresses. This isn’t a thought experiment; it’s a result run against public logs. What it punctures is a security-by-obscurity assumption: encrypting the trace and stuffing it back into the client was supposed to protect IP without the liability of storing it, but the ciphertext’s interchangeability became the handle — good for reconstructing someone else’s hidden reasoning, and for scooping up whatever user credentials were sitting in those logs. All three labs are named, in a rare reproducible first-hand disclosure. (See also the Stolen Thoughts site and the HN thread.)

OpenAI starts putting ads in ChatGPT

OpenAI is testing ads in the Free and Go tiers of ChatGPT, adding the UK, Mexico, Brazil, Japan, and South Korea to earlier pilots in the US, Canada, Australia, and New Zealand; paid tiers (Plus, Pro, Business, Enterprise) stay ad-free. The commitments: ads are clearly labeled, don’t influence the answer itself, advertisers get only aggregate view/click data and no chat content or identity, no ads for under-18s or near sensitive topics like health and politics, and free users can opt for fewer ads in exchange for a smaller daily message allowance.

The thing to watch isn’t “ads in AI” — it’s that the revenue model just set. When a single answer is both the product and the ad slot, “answers aren’t influenced by ads” becomes the load-bearing trust claim, and it’s one you can’t verify from the outside; it rests on the provider’s discipline. For builders, OpenAI’s own framing is the signal: it says ads help fund the infrastructure behind free and low-cost access — serving free users at this scale needs a revenue stream of its own.

Anthropic adds invisible watermarks to Claude’s text

Anthropic signed the EU AI Act’s Article 50(2) Code of Practice on transparency of AI-generated content and will embed a machine-readable, invisible watermark in text from new Claude models launched in the EU on or after August 2, 2026. Per its help center, the text mark is a statistical signal that’s invisible, doesn’t change meaning or readability, survives copy-paste, and may persist through some edits; image files (.png/.jpg/.svg) carry signed C2PA provenance metadata instead. It applies worldwide, not just in the EU, across the API, the Claude apps, Claude Code, and access via AWS, Google Cloud, and Microsoft Foundry. Support for pre-August-2 models is coming, with no date given.

One boundary is written into the official notes: a detected mark means Claude may have processed the content — not that Claude authored it. Run your own writing through Claude to polish it and the output can come back carrying the same mark. That squares with the watermarking survey covered earlier: a watermark can flag “a machine touched this,” but can’t separate “machine-generated” from “machine-assisted,” so treating it as an arbiter of copyright or originality points it at the wrong job.

Two OpenAI exits surface on the same day: the ethics lead and the longtime COO

Two personnel stories broke on the same day. Per the FT, OpenAI’s head of ethics, Chloé Bakalar, quietly left last month, less than a year after joining — she was the company’s only dedicated ethicist, and it reportedly won’t replace her directly (FT). The same day that news surfaced, Brad Lightcap — who joined in 2018, spent four years as CFO, and served as COO from 2022 until moving over to lead special projects earlier this year — announced he’s leaving to “start something new,” staying a few more weeks (Lightcap on X).

Different in kind — one a longtime operator leaving to found something, the other the sole dedicated ethics role vacated and left unfilled — but the Bakalar story is the one with a governance edge. OpenAI’s position, per the FT report, is that AI ethics doesn’t live with one owner or team. From a governance angle, that’s exactly the arrangement where accountability most easily evaporates.

The Gemini app crosses 1 billion monthly users

Google says the Gemini app passed 1 billion monthly active users, the fastest-growing product in company history and Google’s 14th to clear the billion mark (after Search, Gmail, Android, Maps, and others). The usage data is the more interesting part: 63% of users now talk to Gemini directly by voice, it generates more than 150 million images a day, and it has over 100 million active users on iOS.

The scale is the milestone, but the more telling number is the 63% — when nearly two-thirds of users talk to the app by voice, voice has stopped being a side feature, even if that stat alone doesn’t say most interactions happen that way. And 150 million images a day says generative imaging is a high-frequency staple now, not a novelty. The number rides Google’s distribution across its whole ecosystem, so it shouldn’t be read straight against a standalone app’s growth.

Google’s medical AI “AMIE” runs a first real-time video consultation study

Google Research reports that AMIE, its research medical dialogue system, can now do real-time video consultations: reading visual and audio cues at once, doing a “virtual physical exam,” and reasoning toward a diagnosis in real time. In a randomized evaluation using patient actors and primary-care physicians, AMIE scored well on history-taking thoroughness, diagnostic accuracy, management appropriateness, and communication quality, and the patient actors preferred video over text chat.

The status line matters: this is a research system, and Google is explicit that more work is needed before responsible real-world clinical use — don’t read it as “AI video visits are live.” The evaluation used patient actors, not real patients, in simulated consultations, so it speaks to capability ceilings, not to safety in an actual clinic. Read it as a capability advance, not as medical advice.

Mistral bets on “sovereign AI”: in-region inference, open models, homegrown compute

Mistral announced three things at once: Regional Endpoints are now generally available, letting customers pin inference to Europe or the US, plus a Priority Tier with SLAs; the platform will host third-party open models — GLM-5.2 from Z.ai is the first — under the same regional and service guarantees as Mistral’s own; and it’s assembling European firms (ASML, CMA CGM, and others) into a compute coalition targeting up to 1 GW of capacity by 2030. The through-line is “sovereign AI” — keeping data and compute inside European jurisdiction.

Two dimensions to keep apart here: “sovereign/in-region” is about which jurisdiction runs inference and whether you can switch providers, while “open vs. closed weights” is about whether model weights are public — different questions. Mistral’s pitch is mostly the former: optional regional inference plus contractual guarantees to address European buyers’ compliance and data-residency worries. Hosting open-weight models like GLM-5.2 widens the model menu, but whether your data truly “stays in Europe” depends on inference landing on its European infrastructure, not on whether the model is open source.

Empirical study: with military AI decision support, humans distrust more than they blindly follow

Ryan Shandler and colleagues, in the Journal of Conflict Resolution, built a high-fidelity replica of a military targeting decision-support system (DSS) and had 2,015 Israeli military personnel make strike decisions with it across two experiments. The counterintuitive result: the dominant effect isn’t the automation bias everyone fears (blindly trusting the AI) but algorithmic aversion — especially in high-collateral-damage scenarios, people leaned on their own judgment and overrode the AI, and how strongly they did so correlated with their baseline trust in AI. Adding a partial “explainable AI” feature — surfacing the key inputs behind a recommendation — reduced that aversion.

This continues the “humans miss a third of the threats” thread but pulls the conclusion back a step: at least in this simulation, the problem isn’t people deferring to the machine — the machine faces a trust deficit. Note how specific the sample is: Israeli military personnel, simulated rather than live engagements, so the findings may not transfer to other militaries or real battlefields. On the explainability result, the authors’ own read is that surfacing the key inputs encouraged more deliberate evaluation of the recommendations, not blind acceptance. Even so, “making AI more trusted” and “making AI more trustworthy” are different properties, and whether reduced aversion translates into better judgment in the field is the follow-up question this study leaves open.

Stanford HAI: data brokers broadly ignore California’s privacy law

Stanford HAI’s Jennifer King, Daniel Ho, and colleagues studied how 522 registered data brokers comply with California’s 2023 Delete Act, which requires data brokers to register with the state and report annually on how they handle consumer rights requests — the deletion and correction rights themselves come from California’s broader privacy law, the CCPA. Their manual review: only 9% of brokers fully met the reporting requirements, 45% submitted no rights-request metrics at all, and among the 250 brokers whose request processes the team audited, 64% used at least one “dark pattern” design that adds friction to filing a request.

The direct upshot for readers: this bought-and-sold personal data feeds marketing profiles, risk-scoring, and people-search products — the machinery behind the ads and offers you see, and some of the background checks run on you — and there’s a visible gap between privacy law on paper and in practice. Jennifer King’s read is blunt: without clear reporting requirements and consistent financial consequences, businesses simply won’t do it.

One line for today: stuffing the encrypted chain-of-thought back into the client was meant to please both sides — protect the IP, dodge the storage risk — but the ciphertext’s cross-session interchangeability became the biggest opening; security-by-obscurity tends not to hide the thing an attacker most wants.