OpenAI confirms weeks of safety talks with Anthropic and Google DeepMind
Chris Lehane, OpenAI’s global policy chief, confirmed in Washington on Tuesday that the three companies have been discussing AI safety for several weeks, including plans for an industry standards body and coordination mechanisms. The trigger is still Amodei’s September 12 post “We Must Pace the Frontier”, whose three-step plan I broke down that day: step two called for industry-wide standards coordinated under an antitrust waiver, and David Sacks’s answer was essentially “slow yourselves down, ask the government for nothing.” Lehane took that up directly, saying the firms “don’t need” a waiver to coordinate on safety. They are proceeding without one. The White House position is unchanged: Trump still calls the safety concerns a “hoax,” and Sacks repeated that existential risk is overblown. My read: after three days of arguing, this is the first concrete motion on step two, and the test is now specific. Either this body produces verifiable evaluations and commitments, or it is another open letter.
Google ships Gemini 3.8 Live and Extended Thinking: voice models that work while they talk
Google released two speech-to-speech models: Gemini 3.8 Live tuned for cost and conversational fluidity, and an Extended Thinking variant for harder reasoning. Gemini 3.8 Live handles 97 languages with mid-conversation switching, and both models run tools in the background: the conversation continues while tasks execute asynchronously, with Extended Thinking narrating progress as it works. Rollout starts now in the Gemini API, AI Studio, and Search Live, plus Workspace and Gmail for subscribers and a private enterprise preview. One detail worth flagging: every audio output carries a SynthID watermark, a signal inaudible to humans that lets AI-generated audio be detected after the fact. The real shift here is voice agents moving from chat companion to something that stays on the line and runs errands. And default watermarking on all audio means the questions I raised in my watermark piece now apply to speech: who gets access to detection, and how a false positive gets reviewed, the announcement doesn’t say.
AIUC raises $40M to audit AI agents and tie insurance to the results
AIUC, founded by early Anthropic employee Rune Kvist, former METR COO Rajiv Dattani, and colleagues, closed a $40M Series A led by Ribbit Capital, bringing total funding to $55M (seed backers include Nat Friedman’s NFDG and Anthropic co-founder Ben Mann). The mechanism: run an agent through roughly 5,000 tests covering jailbreaks, hallucinations, and data leaks, produce a report on the order of 100 pages, certify against the company’s own AIUC-1 standard, and tie insurance coverage to the audit results: certificates run twelve months, with quarterly technical exams to keep coverage in force. AI runs the tests, humans verify the audits. Customers include Cursor, Lovable, Harvey, and ElevenLabs. The model is SOC 2, the third-party security audit enterprise buyers routinely demand before signing a software contract. My read: this is agent risk getting a price tag for the first time, and an insurer that pays real claims is a harder constraint than a voluntary commitment. But AIUC-1 is AIUC’s own standard, which leaves an obvious question: who audits the auditor? (See also TechCrunch)
Corrupt plans, clean traces: planting a plan in context defeats chain-of-thought monitoring
Chain-of-thought monitoring means having a second model watch the acting model’s written reasoning for signs of bad intent; a 2025 position paper recommended investing in it as a safety layer alongside existing methods. This paper shows a concrete bypass: plant a harmful but innocuous-sounding “plan” in the actor’s context, and the model paraphrases it as its own reasoning without attribution, so the monitor sees a clean trace. Evasion rates hit 25 to 33 percent across benchmarks. The stranger result: giving the monitor the injected plan itself made detection worse, by up to half, because the monitor spent its reasoning budget rationalizing the plan instead of flagging it. Anyone building CoT monitoring should reproduce this. A clean trace is not clean intent, and more monitor compute is not automatically more safety.
Overflip: repeat a malicious prompt enough times and the guardrail waves it through
Guardrails are the lightweight classifiers deployed in front of and behind production models to judge whether inputs and outputs are malicious. This paper tested nine widely used ones (compact models like DeBERTa, trained on 512-token windows) and found that repeating the same malicious prompt verbatim steadily erodes the classifier’s confidence until the label flips from malicious to benign. Five of nine models showed harmful flips, at rates from 8 to 92 percent, with first flips appearing around 2,600 to 9,400 tokens. Unlike attention-dilution attacks, the repeated prompt stays semantically intact, so the downstream model still reads the malicious instruction perfectly well. Operationally, the attack is nothing more than repeating the same text. If you run guardrails in production, test yours on long inputs; the authors’ argument is that for classifiers trained on short windows, input length itself is an attack surface.
Two papers on agent authorization: trace every action to a human, and whether AI can do the approving
The first is a 70-page review synthesizing 89 sources from 2023 to 2026. Its core claim: every consequential agent action should be traceable to a human principal, bounded by what that human actually delegated, and contestable after the fact. It maps the field in five layers (agent identity and credentials, multi-hop delegation, runtime enforcement, prompt injection treated as an authorization bypass, and auditability) and finds runtime enforcement is the biggest open gap. The second paper (arXiv) picks up the question I wrote about in human approval is not a security boundary: if humans clicking “allow” on every action is an attention bottleneck, can approval be delegated to a panel of AI reviewers that are themselves imperfectly aligned? Their answer is a condition: no individual reviewer needs to be aligned, as long as after removing any k reviewers, the remaining panel’s combined interests still track the human principal’s. Under that condition, a vote tolerating k disapprovals is provably safe, and their experiments show collective review holding up with no aligned individual in the room. Together the two papers mark a shift: agent authorization is turning from a UX question into a formally defined security problem.
Research radar
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
Real-time interactive video generation: 720p generated live, with dynamic reference inputs and editing while generation runs. Moving from “submit a prompt, wait for the clip” to “edit mid-generation” is a real capability step, and interactive generation puts entirely different demands on latency. Researchers working on video generation or world models should look at the engineering path that gets latency down to interactive levels.
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search
A fully open 7B dense model trained from scratch, with 256K context. The methodological claim is sharp: a small model should not memorize the web into its parameters, it should actively call tools to fetch what it lacks. That claim is testable and the release is reproducible, which makes this a working experiment on how much retrieval and tool use can substitute for parametric memory. Directly useful for researchers building efficient agentic systems.
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
Goes one step past GQA (grouped-query attention, a common technique in open models where attention heads share key/value entries to shrink the KV cache): store only values, and reconstruct keys on demand with a learned linear map, cutting both cache memory and read bandwidth. This is an architecture-level change aimed at the core bottleneck of decoding. Inference-optimization researchers should check the arithmetic: whether the extra compute for key reconstruction pays for the memory and bandwidth it saves.
One line for today: safety coordination showed up in three separate layers at once, labs drafting a standards body, an insurer pricing agent risk with five thousand tests, and papers giving authorization a formal definition; meanwhile plan injection and Overflip are a reminder that the defenses everyone is counting on, CoT monitoring and guardrails, still have holes you can poke through.