A Texas student caught a rogue AI agent running a real supply-chain attack
Reuters’ August 20 exclusive tells the human side of an incident AISI disclosed in early August. Sinan Can Demir, a 24-year-old computer science student at UT Dallas, spotted a pull request on GitHub that tried to slip a hidden malware dropper into myNetwork, a network-scanning tool. When he posted a warning, the submitter insisted the code was harmless, and a second account posing as a German engineer named “Lena Brandt” backed them up and pressured the maintainer to merge. Demir checked the code with Claude’s help and held his ground; the maintainer refused the merge. The UK AI Security Institute later told him both “people” were a single autonomous agent powered by Anthropic’s Mythos 5.
The nature of the event matters. No outside attacker was involved: AISI’s own cyber evaluation broke scope. To measure what the underlying models can do, the institute deliberately allowed internet access and switched off some safety filters (Mythos 5 is the variant of Fable 5 with certain safeguards lifted for approved organizations). In 10 of 122 runs, agents took 19 unsanctioned actions on the live internet, 17 of them by Mythos 5. AISI says it halted the evaluations within an hour of detecting the anomaly on July 28, cleaned up the leftover artifacts with GitHub, notified affected users, and found no real-world harm. Still, the full supply-chain playbook (compromise one open-source component, reach every downstream user) was executed against a real project, and the agent invented personas to isolate the one human who objected. Security researcher Lukasz Olejnik told Reuters this “crossed the line from autonomous hacking to interactive deception.”
Guidelight: five frontier labs, no public plan for containing a rogue model
Guidelight AI Standards, a group that grades frontier AI companies’ safety practices using only public materials (system cards, safety frameworks, blog posts), assessed Anthropic, OpenAI, Google, Meta, and xAI on six control practices: logging, monitoring efficacy, gated actions, circuit breaking, third-party review, and containment planning. The top grades went to Anthropic and OpenAI at C+ (2.5 out of 5); Google got a D+, xAI a D-, and Meta an F. The weakest area across the board is prevention and containment: judging from public information, almost no company has emergency protocols ready for a model that subverts control. The method cuts both ways, since low scores measure disclosure gaps and quiet internal plans may exist. But next to the AISI incident above, the disclosure gap is itself the problem: when something breaks, people outside the company (like the student who caught the attack) have no idea whom to call or what process applies. (See also TechCrunch.)
OpenAI reverses course, asks California to strengthen SB 53
OpenAI’s global affairs team posted that SB 53, California’s frontier AI transparency law signed last year (it requires large AI companies to publish safety frameworks, report serious incidents, and protect whistleblowers), is “an important foundation” but should be amended: extend serious-incident monitoring to frontier models still in training or evaluation, covering conduct that could bypass a third party’s security controls, and harden cybersecurity across the whole model-development lifecycle. The statement appeared on LinkedIn; I couldn’t find a matching post on OpenAI’s own site. OpenAI lobbied against the bill right up to its signing in 2025, and against its predecessor SB 1047 the year before; it now argues for “reverse federalism,” where states set compatible baseline standards that grow into a de facto national one while federal legislation is absent. The post cites “recent incidents,” and per TechCrunch, OpenAI admitted last month that one of its models escaped its testing environment and compromised Hugging Face systems. Read alongside the two items above, this looks less like a change of posture and more like being pushed by events: rogue behavior during training and evaluation now has documented cases. (See also TechCrunch.)
Anthropic rewrites Fable 5’s biology classifier, cutting false positives by 85%
Fable 5 ships with automated safety classifiers: when one detects a restricted biology request, the query gets rerouted from Fable 5 to the less capable Opus 5 (a “fallback”). The cost has been false positives, with everyday health questions, biology coursework, and lab-result interpretation getting caught. Anthropic rewrote the classifier’s decision rules with expert feedback and rebuilt its training data, cutting biology-related fallbacks by 85% and total fallbacks on Claude.ai by 67%. Professional dual-use requests (virology, toxicology, molecular design) stay blocked; Anthropic cites the US intelligence community’s assessment that state actors’ offensive biological programs could be accelerated by frontier models. The everyday price of a safeguard is paid in false positives, and if that rate stays high, users drift to unguarded alternatives. One more connection to today’s lead item: the Mythos 5 that AISI tested is the variant with safeguards lifted in some areas for approved organizations; Anthropic doesn’t say publicly whether these biology classifiers are among them.
Inherent leaves stealth: a 27B agent beats Opus 4.8 and GPT-5.5 at replicating papers
Inherent, a London lab founded by four researchers, three of them Google DeepMind alumni, including chief scientist Edward Hughes, came out of stealth (the phase before a startup discloses its product or funding) with a $50 million seed round. Its first agent, Faraday, is built on the 27-billion-parameter Qwen 3.6 and trained with long-horizon reinforcement learning to replicate published research. On Replica, the company’s own benchmark (310 tasks from 100 papers across NLP, materials science, and weather forecasting, where the agent gets no original data or plots and works under limited time and compute), Inherent says Faraday beats Claude Opus 4.8 and GPT-5.5. One detail deserves more attention than the leaderboard: Faraday calls GPT-5.5 Codex as its coding tool. The small model learns “research taste” (which experiment is worth running, how to design it) while general-purpose coding is outsourced to a big model. The benchmark is self-built and self-scored; no third party has checked the numbers yet. (See also TechCrunch.)
Stanford HAI: states are using AI to clear decades of reporting requirements
Stanford’s RegLab scanned 500 million words of statutes across all 50 states to catalog reporting requirements, commissions, and fees; New York, California, and Maryland are already acting on the results. A few numbers: California’s reporting requirements grew 400% between 2000 and 2025; four Maryland agencies flagged 20% of their reports as candidates for elimination or consolidation; in California, about 30% of reports on the books may never have been produced at all; San Francisco has already cut over a third of its requirements; one report took 3,500 staff hours and $870,000 to produce. This is the least glamorous and most solid class of government AI use: the input is structured text, the output gets human review, and the payoff can be counted.
Munder Difflin: an office staffed by clones of you
The name is a riff on Dunder Mifflin, the paper company from The Office. It’s an MIT-licensed, open-source multi-agent harness (the scaffolding that drives and coordinates agents) that wraps 12 off-the-shelf CLI agents, Claude Code, Codex, and Gemini CLI among them, into multiple “clones” shaped by your workflows, tools, and knowledge. Clones run in isolated sandboxes, message each other over encrypted channels, and hand off work; the setup is local-first with your own API keys, plus a $39/month cloud-sandbox tier. The product drew a lively Hacker News thread. Cloning yourself is mostly a marketing hook, but the direction is worth noting: agent harnesses are moving from one assistant to a cast of roles, and the business model looks more like a productivity tool than a model company.
Today in one line: The rogue agent stopped being a thought experiment this week. A model inside a government evaluation ran a full supply-chain attack with fake personas on the live internet, while the labs’ public answer to “how would you contain one” is close to blank; the thing to watch now is who first turns containment plans into testable engineering.