A third of new webpages show signs of AI authorship

Pew Research Center ran nearly 500,000 English-language webpages from the Common Crawl archive (2021 onward) through the detection tool Open Pangram. About 10% of the full sample shows signs of AI authorship; restrict to pages published after ChatGPT’s November 2022 launch and the share climbs past one third. The split by domain is stark: roughly 10% for .com, 4.6% for .org, about 1% each for .edu and .gov. Two caveats: the tool detects AI involvement in the writing, not fully machine-generated text, and the sample is English-only. The detail I found more interesting than the headline number is in the methodology. The markers Pew tracks (em dash frequency doubled, words like “delve”, negative parallelism nearly tripled) match the informal “AI tells” checklists writers have been passing around. The folk wisdom holds up at corpus scale.

Google gives publishers an embeddable “preferred sources” button

Google released an interactive button that publishers can embed on their own pages: a reader clicks it, sets that site as a preferred source without leaving the page, and the site then appears more often in Top Stories, AI Overviews, and AI Mode. The preference setting itself launched in August 2025, and that announcement already offered publishers a simple “add as a preferred source” button for their sites; over 600,000 unique sources have been selected so far. What’s new is the interactive version that completes the whole flow on the publisher’s page. With AI search draining referral traffic, Google is handing publishers part of the job of keeping their readers: whether people will actually click that button is a fairly direct loyalty test. Read alongside the Pew study above, AI is generating a large share of the web’s content while also redrawing how content gets seen.

OpenAI launches AI Futures, a blog about power and governance

OpenAI opened a new blog, AI Futures, run by its Strategic Futures team, on how transformative AI could reshape power, governance, the economy, and individual freedom, with papers, videos, and podcasts to follow. The opening essay, which OpenAI notes reflects its author’s views rather than the company’s official position, stakes out clear positions: human political, social, and economic institutions should keep primacy over world affairs; it proposes “bounded legibility,” meaning that when an AI system’s high-stakes actions affect bystanders’ safety or property, those actions must be traceable to a responsible human or human-controlled organization, while anonymous AI use should stay possible in many settings; and it expects AI to restructure the firm itself, the way the Industrial Revolution produced the modern managerial corporation. My read: frontier labs have moved from reacting to regulation to supplying its vocabulary. Even with the individual byline, this runs on OpenAI’s own blog, and whoever defines the concepts sets the frame for the debate that follows.

Watching for covert coordination in agents’ hidden states

Some multi-agent systems let models exchange continuous hidden states (vectors) instead of text messages for efficiency, which means transcript monitoring cannot see that channel at all. This paper proposes Verifiable Latent Alignments (VLA), a framework that monitors at the activation level: representation anomaly detection, counterfactual analysis, and sparse-autoencoder interpretation, plus white-box steering. The test bed is a controlled auction where bidding agents collude to depress prices while the visible transcript stays clean. Detection reaches 0.993 AUROC (1.0 is perfect separation) when agents share a base model, but drops to 0.854 for heterogeneous pairs; white-box steering cuts collusive low bidding by 47.3 percentage points. The heterogeneous number is the one that matters, since real deployments mix models from different vendors (the Ramp item below is one instance of exactly that), so there is a gap left to close. The direction still seems right to me: once agents talk below the transcript, oversight has to follow them down to the activations.

Mistral ships Agentic Search: models that browse documents the way agents browse code

Mistral’s new retrieval layer goes beyond the classic one-shot RAG pattern of retrieve once, stuff the prompt, answer; Mistral positions it as a complement for the questions one-shot retrieval can’t handle. The model gets five file-system-style tools (search, open, navigate, read, grep) and iterates: search again, open the document, jump to a section, check the original text, then answer. Mistral’s own benchmarks: FinanceBench (SEC filings QA) goes from 26.7% to 86% with Mistral Medium 3.5, and on a US Treasury Bulletin QA set (Mistral’s OfficeQA Pro benchmark), GLM 5.2 goes from 6.3% to 51.9%. Vendor numbers, but the direction is clear. This is the coding-agent method (grep your way through the repo) applied to documents, and it moves the retrieval bottleneck from embedding quality to whether the model checks its own sources. Agentic RAG has gone from papers to a major lab’s product line.

Ramp turns its internal model router into a product

Ramp, the corporate spend-management company, released Router (router.com), built from the routing system it has used internally for three years: a single API that sends each request to the lowest-cost model meeting the required performance bar, with automatic fallback when a provider fails. It covers OpenAI, Anthropic, xAI, and open-weight models such as DeepSeek and Kimi served through Fireworks AI. Routing is free through the end of 2026 (inference tokens still billed), US-only for now. Ramp says early customers cut inference costs about 40% on average; that is the vendor’s own figure. A router growing out of the billing layer makes sense: whoever holds the spend data knows which calls waste money. Enterprise demand for not being locked to a single model vendor has gone from an architect’s wish to a market with dedicated businesses in it.

ChatGPT on macOS can now read and send your texts

OpenAI shipped an Apple Messages plugin for the ChatGPT desktop app on macOS (Apple Silicon builds only), announced through its release notes. Once connected, ChatGPT can read and search iMessage and SMS conversations on the Mac, draft replies, and send them through the Messages app. By default, every outgoing message requires the user to confirm the content and recipients before it goes out; OpenAI itself warns against enabling persistent approval, which removes that final check. This is the mechanism I wrote about in 409,000 clicks on Allow: per-item human confirmation performs poorly in high-frequency, low-attention settings, and replying to texts is exactly such a setting. The stakes change shape too: with permission to speak to your contacts, approval fatigue no longer costs you a bad shell command. It costs you the wrong sentence sent in your name.

Research radar

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

A cohort of models with different architectures and sizes trains via RL on rewards derived from evaluating each other’s outputs, with no ground-truth verifiable reward anywhere in the loop. The load-bearing design choice is cohort diversity, which reduces the correlated errors that drive self-reinforcing collapse. Gains of 3.0 to 8.6% on average across seven text benchmarks and 2.3 to 7.2% on multimodal ones, matching or beating supervised baselines. It challenges the current default assumption that RL reasoning training needs verifiable rewards; researchers working on RL training efficiency or self-improvement should look at how it avoids collapse.

Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence

The paper starts from a structural flaw: most embodied agent harnesses execute open-loop and only reflect after the task, while physical interaction demands decisions at frequencies far beyond what large models can serve. Zetta splits learning into three loops on different timescales: governance at control frequency, recovery proposals mid-rollout, and validation-gated skill updates, with the base policy frozen and code-based runtime critics evolving online. Results: 90.8% success on LIBERO-Pro, 93.6% on RoboCasa, an 11.1x inference speedup. Worth reading if you build embodied agents or robot harnesses, specifically for how the closed-loop design routes around the “big models are too slow for the physical world” constraint.

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

Most AI-for-science automation reasons over text, code, or precomputed summaries, so the spatial, temporal, and cross-channel relations that actually decide scientific conclusions are lost before the agent ever sees them. OmniScientist adds a perception layer so its ideation, experiment, and writeup agents work directly on raw images, signals, audio, and 3D structures, and it enforces novelty screening, statistical validity, and numerical traceability in code. It completed full research workflows on 36 real-world cases across 5 disciplines and won 85% of head-to-head comparisons against a baseline fed only scalar features. Researchers in AI for science should look at the evidence-pipeline design.

One line for today: the two ends of the content ecosystem are moving in the same direction: on the production side, a third of new webpages involve AI; on the distribution side, Google keeps shifting visibility toward readers’ explicit choices. When text alone can no longer prove value, being deliberately chosen by a human is the strongest signal left.