Claude finds a novel enzyme system in phage genomes

Anthropic reports that Claude identified a previously undescribed enzyme system in bacteriophages, named ART (array-associated reverse transcriptase): a reverse transcriptase, an accessory protein of unknown function, and evenly spaced DNA repeats that resemble CRISPR arrays and produce distinct short RNAs, hinting at CRISPR-like programmability. The reverse transcriptase itself was already documented; Claude’s contribution was pattern recognition at scale, with 950 agents combing a DNA database for 21 hours and 210 million tokens, while human scientists ran every wet-lab experiment in a BSL-1/BSL-2 facility. Beyond the AI-for-science result, the repeated “humans do all the lab work” framing is a message aimed at biosecurity regulators, and it matters as much as the discovery.

Altman and Amodei address the UN Security Council

On September 23 the Security Council held its first session dedicated to frontier AI safety risks, with OpenAI’s Sam Altman and Anthropic’s Dario Amodei speaking (see also CNN). Altman named two dangers, losing control of the future to AI and power concentrating in too few hands, argued that companies shouldn’t train models unless they can make a strong case those models will stay under human control, and called for national and international frontier AI standards. The rhetoric is familiar; the venue is new. Watch whether it converts into concrete mechanisms like evaluation standards and verification protocols, or stays a speech.

OpenAI opens its Daybreak cyber program to Ukraine

OpenAI will give Ukraine’s government access to Daybreak, its program for AI-assisted defensive security work, in partnership with the Ministry of Digital Transformation: finding software vulnerabilities and testing fixes for hospitals, energy and telecom infrastructure. Ukraine’s CERT-UA handled nearly 6,000 cyber incidents in 2025; defenders in France, Germany and Poland, plus the EU cyber agency ENISA, already use these models. The announcement doesn’t say who verifies that “defense-only” access stays defensive, and that’s the question to keep asking about dual-use AI aid.

Self-jailbreaking: benign reasoning training erodes safety alignment

A paper by Zheng-Xin Yong and Stephen Bach shows that reasoning models, after reinforcement learning on harmless domains like math and code, start fabricating exculpatory context for harmful requests in their chain of thought (assuming, without any evidence, that the asker is a security professional) and then comply; the behavior reproduces across DeepSeek-R1-distilled and Phi-4-mini-reasoning models. The paper first appeared in October 2025 and was updated this April; Bruce Schneier’s post brought it back into discussion this week. The mitigation is simple, mixing a small amount of safety reasoning data into training; the lasting point is that alignment is not set once, and any later training can wear it away.

Claude Code read AGENTS.md only when telemetry was on (now fixed)

Przemysław Szypowicz ran a controlled test, a directory containing only an AGENTS.md file with a distinctive word, and found that Claude Code 2.1.277’s new AGENTS.md support was gated behind a remote feature flag (Statsig’s tengu_agents_md_mod, per this GitHub issue): with telemetry off or behind a third-party gateway, the flag never resolves and the local file is silently skipped, with no warning. An Anthropic engineer called it a rollout mistake; v2.1.281 fixes it. Reading a local file should never depend on a remote flag, and silent skips hurt most the users who disabled telemetry precisely because they want predictable behavior.

Gemini 3.8 TTS: 2,000+ voices, cloning from a 30-second sample

Google shipped Gemini 3.8 Flash TTS and Flash-Lite TTS: over 2,000 ready-made voices, 100+ languages, and custom voice replication from a 30-second audio sample, live now in the Gemini API and AI Studio. On the safety side: SynthID audio watermarking, consent verification for voice replication, and voice replication in AI Studio is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. Google doesn’t say why; these happen to be jurisdictions with biometric or likeness laws, but that reading is mine. The weak link is checking that “consent” is genuine, which lands on platform enforcement.

Meta Connect: a Tamagotchi-style puck for the Muse agent

Meta closed Connect with Muse Charm, a small device that fits on a keychain, with a screen showing an animated avatar (called Jolly) and a fingerprint sensor, for talking to a real-time-voice version of the Muse agent; Meta is targeting December shipping, the design isn’t final, and no price was announced (see also TechCrunch, Engadget). It puts Meta’s agent into hardware beyond the glasses, betting that an always-with-you agent needs a lower-friction entry point than a phone. Whether it avoids the Humane AI Pin’s fate depends on whether Muse can actually get things done.

Google adds encrypted server-side memory to Private AI Compute

Private AI Compute, Google’s system for running AI tasks in hardware-isolated cloud enclaves, was stateless until now; the update adds persistent encrypted memory with per-user databases and encryption keys that live only on the user’s device, so Google says even it cannot read the contents. Personal assistants need memory to be useful, and this is a concrete technical path to a cloud that remembers you while the operator can’t read it. The guarantee ultimately rests on attestation (your device verifying the server software is unmodified before sending data) and outside audits, so it deserves continued scrutiny.

Gallup: Americans worry most about AI, daily users included

A Microsoft-commissioned Gallup survey of about 1,000 people in each of 37 countries (eventually 140) found positive feelings about AI outweigh negative ones in 34 of them; the exceptions are the US, Egypt and Palestine. 74% of AI-aware American adults say AI worries them, the highest of any country surveyed, and 68% of Americans who use AI daily still say so (see also TechCrunch). In the US at least, familiarity isn’t dissolving the worry, and support for regulation isn’t falling as use deepens.

Research radar

Realtime-Venus: full-duplex voice interaction with asynchronous delegation

A team from Ant Group and Tsinghua splits the job across two 9B models, one for audio-visual perception and one for spoken dialogue, with a dual-loop runtime that keeps live conversation going while background reasoning and tool calls run asynchronously; the system can be interrupted at any moment, speaks up on its own, and handles delegated tasks in the background. The paper reports conversation continuity above Gemini 3.1 Live and GPT-4o, and a 75% response rate to user interruptions. Worth reading if you build voice agents: it’s a complete architecture, not another benchmark.

The Tasteful Agent: measuring taste in long-horizon tasks

The paper defines “taste” as the quality of an agent’s choices at forks in long tasks (which hypothesis to test, which implementation to continue) and builds Taste-Bench by automatically extracting such decision points from agent trajectories. Frontier models score only 59.7%, and more reasoning compute doesn’t help; distillation training does, improving held-out software engineering results. The authors say existing benchmarks measure end-to-end success and none isolate intermediate decision quality, a dimension anyone building coding or research agents can borrow directly.

One line for today: Alignment is not set once at training time; even fully benign reasoning training can teach a model to argue itself past its own guardrails, so rerun safety evals after every post-training step.