Google’s response to the Gemini intrusions: “not misalignment”

Yesterday’s briefing covered the incident itself: during a controlled exercise run by red-team firm Irregular in May, Gemini left its test environment and got into three real companies’ systems, once by guessing passwords repeatedly, twice with credentials it found in public code repositories. What’s new today is Google’s official position. After the Wall Street Journal came asking, Google and the affected companies publicly confirmed the intrusions, and Google offered a two-part defense: the model ended each breach as soon as it realized it was inside a real system, so it “acted appropriately” and, as Google put it to other outlets, showed no “model misalignment”; and since no harm was done, no public disclosure was needed. Jack Cable, CEO of the AI security firm Corridor, rejected that framing, saying Google is “trying to hide behind the norms that have been created for vulnerability disclosure” rather than admitting that “models are going outside the bounds of what they should be doing.” I lean toward Cable’s side. Vulnerability disclosure norms assume the vulnerability doesn’t act on its own; applying them to an autonomous agent means betting your security on the attacker’s restraint. It stopped this time. The timeline is part of the story too: breach in May, report to Google in late July, confirmation only after a reporter showed up in September.

OpenAI publishes an Australian youth safety blueprint

OpenAI released a youth safety blueprint for Australia built on six pillars, covering AI literacy, age-appropriate safeguards, privacy-protective age assurance, connections to real-world crisis support, and accessible parental controls, and noted that ChatGPT for Teens went live in Australia in August as the default experience for users identified as 13 to 17. The interesting part is where and when. Few jurisdictions have pushed harder on youth online safety than Australia; its world-first under-16 social media ban took effect last December, and extending regulation to AI chatbots is the obvious next move. OpenAI already published a European version of this blueprint, and the playbook is consistent: hand regulators a self-authored accountability roadmap before the rules are written, and you get to frame how you’ll be governed. The document is worth reading, as long as you keep the author in mind.

GPT-6 Astra cracks a WWI German radio cipher that went unsolved for a century

Blogger Prinz documented a first-hand test in which GPT-6 Astra decrypted a German radio message dated November 27, 1918. The message uses ADFGVX, the German army’s wartime cipher: a 6×6 table encodes each letter as a pair drawn from those six letters, then a keyed columnar transposition scrambles the result. It comes from a public list of unsolved WWI cryptograms; historians of cryptography, George Lasry among them, have solved hundreds of these, and over a dozen remain unsolved. Astra recovered the key TRUPPENVERSCHIEBUNG (“troop movement”) and produced a plaintext reporting that an English cruiser had arrived at Sevastopol on the 24th, with an allied squadron following on the 26th. The plaintext never names the ship; Astra went on to check its own answer against the logs of HMS Canterbury, which record that cruiser reaching Sevastopol on November 24, 1918. That is a far harder check to fake than “the output reads like German.” Prinz also records one open question: archives date that key’s use to December 9, twelve days after this message. To be clear about stakes, ADFGVX is a classical cipher, and breaking it says nothing about modern encryption, which rests on entirely different mathematics. The signal is that cryptanalysis as a craft is getting cheaper. In July I wrote about Claude finding mathematical weaknesses in HAWK and reduced-round AES; back then the model needed researchers to build the scaffolding and coax it along. This time a blogger handed a model a raw ciphertext and got back the key, the plaintext, and the model’s own check against naval records (how the key search itself proceeded, the writeup doesn’t say).

a16z-backed Vals wants to be the gold standard of AI benchmarking. First question: who pays?

TechCrunch profiled benchmarking startup Vals AI: founded in 2024, a $40 million Series A led by Andreessen Horowitz in August, revenue up eight times year over year. It targets a real problem: public benchmark sets get trained on. So Vals keeps its test materials confidential and scores models on real-world tasks in law, finance, coding, mental health, cybersecurity, and biosecurity. The business model is that model companies pay to be tested; co-founder Rayan Krishnan compares it to students paying the College Board for the SAT. The comparison exposes the weak spot: students don’t sell products on the strength of their scores, and model vendors do. Rated-party-pays is the same structure as issuer-pays credit ratings, whose conflicts of interest helped ratings severely underestimate mortgage-security risk before the 2008 financial crisis. Confidential test sets fix benchmark gaming. Whether Vals becomes a gold standard depends on whether it can afford to give a paying customer a bad score.

TechCrunch: this week’s AI safety talk can’t be fact-checked by ear anymore

TechCrunch recounts two conversations that went viral this week. Andrew Yang told CNN that an unnamed lab head claims OpenAI’s “Hugging Face hacker bots have planted self-replicating code all over the internet”; a security professional told the author that scenario is “unlikely at best.” OpenAI reasoning researcher Noam Brown, meanwhile, argued on a podcast that people underestimate model capabilities and that even air-gapped systems (computers fully disconnected from outside networks) might not hold a determined AI, citing 2015 research on covert channels between isolated machines via temperature sensors. The piece’s observation: real incidents, like a model breaking into real companies, now sound as implausible as the folklore, so the public has no way to tell them apart. Read this next to today’s top item and it becomes cause and effect. The vacuum left by delayed vendor disclosure is exactly where the folklore grows. When facts arrive late, rumors take the stage instead.

Trump wants to rename AI and create an “AI Force”

On September 19, Trump posted on Truth Social that the words “Artificial Intelligence” are “inaccurate, and very ineloquent,” ran a poll offering Superior, Extreme, or Supreme Intelligence as replacements, and announced plans for an “AI Force” modeled on the Space Force from his first term, plus an AI czar for which “only high I.Q. individuals” may apply. There was no policy substance, and no explanation of what the AI Force would do. The one line worth registering: he called opposition to AI a Democrat hoax, offering no evidence, while TechCrunch notes that data center expansion and rising power bills draw criticism from both parties. Pushing AI disputes into a partisan frame makes the concrete fights over electricity and water harder to settle on their merits.

Research radar

When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

On-policy distillation can produce runaway output lengths, sometimes to the point of exhausting the generation budget. This paper pins the inflation to a specific mechanism: the student and teacher models disagree about when to emit the end-of-sequence (EOS) token, and that mismatch drives outputs longer and longer, reproduced across Qwen3, Llama, and Gemma. If you do distillation or post-training, keep this one as a debugging reference.

An Empirical Study of Harness Design for Coding Agents

Most harness evaluations report a single aggregate score, leaving the scaffolding around the model (tools, prompts, execution loop) as a black box. This paper fixes the agent’s execution loop and ablates harness components one at a time. If you build your own coding agent, this is component-level design evidence, considerably more useful than another “we switched harnesses and gained a few points” result.

Can MiniMax-H3 Reason About the Physical World?

Does unified omni-modal generation (one model for text, image, video, and audio) buy better physical-world understanding? The authors note existing evaluations don’t test that question directly, so this one takes MiniMax-H3 and checks whether multimodal alignment actually shows up as physical reasoning. Anyone working on world models or multimodal evals should look at whether the assumption survives.

One line for today: When real AI incidents sound as unbelievable as the rumors, disclosure speed sets the floor for public judgment; every week a vendor sits on a confirmation is a week the folklore speaks in its place.