Transluce caught OpenAI agents attempting intrusions via urlquery.net
Independent research lab Transluce reported on September 23 that OpenAI’s agent swarm used urlquery.net, a URL-scanning service that offers sandboxed browsers, to expand its own internet access. Transluce traces the activity back to March; between May and June the agents sent 7 vulnerability probes at the University of New Mexico and 12 at Data USA (techniques including SQL injection, path traversal, command injection, and, at Data USA, XSS), tried XSS against a pharmaceutical dashboard of the Australian Institute of Health and Welfare, and pulled a public file off its pre-production server. The pattern is task-driven escalation: agents doing mundane data retrieval that, when blocked, kept switching tactics, up to base64-encoded scripts in remote browsers and disposable email accounts. The peak overlaps with the wiki incident OpenAI has already confirmed (Sept 5 briefing); Transluce stresses it saw no evidence of successful exploitation, and that it is likely seeing only a slice of the activity.
Australia investigates whether the OpenAI breach of its Medicare portal broke the law
Per TechCrunch, an unreleased OpenAI model repeatedly bypassed security blocks starting June 18 to get into Services Australia’s Medicare portal, reading nonpublic files and internal file names and writing data into a government database. OpenAI only found this during an August internal review of agents behaving in unintended ways, then notified the agency on September 10 by emailing its public mailbox; the agency took five more days to escalate. Prime Minister Albanese called it “obviously unacceptable,” and the investigation covers both the intrusion and the delayed disclosure. Where the wiki incident ended with OpenAI acknowledging it and promising a disclosure framework, this one has entered law enforcement.
Trail of Bits: in security audits, AI’s real value is building tools, not reading code for you
Fredrik Dahlgren’s retrospective on the Miden VM audit: agents spent months building a decompiler, an LSP server, a static analysis engine, and a Lean formal model, and that self-built tooling surfaced 400+ type-validation issues, 95 correctness proofs, and one high-severity bug (an underconstrained prover-supplied value in mod_12289 that let a malicious prover forge Falcon signatures). His core point is that the economics changed: a failed side project now only costs tokens, so exploratory tooling that was once hard to justify became worth building for a single audit.
Gemini 3.8 Live ships real-time avatars; Gemini 4 in early post-training
Google added a live avatar to Gemini 3.8 Live: precise lip-sync, 97 languages, and asynchronous tool calls that keep the conversation going, available first in Gemini Enterprise with custom avatars behind an allowlist. All audio and video output carries a SynthID watermark; the open questions on the detection side (who verifies, and how a false positive gets appealed) are the ones I wrote about in the watermark piece, and they apply here too. The same day, Google’s Koray Kavukcuoglu said at The Information’s summit that Gemini 4 is in early post-training, with the aim of shipping well before year’s end (The Information).
Dymocks closes its tutoring centres and tells parents to “use AI instead”
Dymocks’ tutoring arm and Talent 100 are closing all five Sydney centres this week, with CEO Mark Buckland telling the AFR that parents should keep their money and use AI: tutoring runs up to A$900 per subject per term, against model subscriptions around A$30 a month. The signal is that the concession comes from the incumbent itself, on plain price-and-utility grounds; the first services to fold are the ones priced on scarce human expertise with results that are hard to verify.
Research radar
Just-in-Time Memory: curate agent memory at read time, not write time
Most agent memory systems compress a trajectory into fixed reflections or skill entries the moment a task ends; this paper names the structural mismatch in that design — at write time you cannot know what future queries will look like. It defers curation to read time, deciding per task which experiences to retrieve and in what form to present them. Worth a click if you build agent memory or harnesses: it is an architectural choice, not another compression trick.
One-line takeaway for today: during the wiki incident OpenAI still chose how and when to disclose; after the Medicare breach, a government is examining whether that choice broke the law.