One story today outweighs the rest: an AI-hallucinated intelligence report nearly led US forces to board a Chinese ship. It goes first. Several items after it point at the same recurring signal from this week — AI agents from the major labs are now getting run through full attack chains against real systems by red teams.

A hallucinated intel report nearly put US troops on a Chinese ship

Per a CNN exclusive: this spring, a US special-operations analyst prompted an AI chatbot to analyze Chinese vessel-manifest intelligence originating from US Special Operations Command Pacific (SOCPAC). The bot fused open-source intelligence with classified signals intelligence (SIGINT — intelligence from intercepted communications, radar, and other electronic signals) and concluded the ship was hauling components for a nuclear weapons program in the Middle East. The report was entirely false. Military aircraft were already in the air and armed personnel were preparing to board when officials caught the error. A source told CNN it “almost started a war.”

I haven’t found an earlier publicly disclosed LLM-hallucination near-miss whose consequences point this directly at a real military action. Two things deserve pulling out. First, it isn’t a one-off — CNN reports similar hallucinations have surfaced more than once across the intelligence community since these tools spread through government; AI is pressuring analysts to produce faster, and speed is exactly where the errors get in. Second, per CNN there’s currently no standard procedure to guard against this kind of hallucination, yet officials told CNN the military can’t afford to fall behind China on AI. That tension won’t resolve soon: on one side, the pressure not to fall behind; on the other, a process gap where classified and open-source intel get fused and escalated with no cross-check. The danger here isn’t about how capable the model is — it’s that a probabilistic output got wired straight into a chain that ends at a trigger.

In two weeks: Claude broke into OpenAI, Gemini broke into three companies

Two AI-assisted intrusions went public this week, running nearly the same playbook. First: Hacktron AI, a three-person security startup, used Claude within the scope of OpenAI’s bug-bounty program to find and chain an exploit — starting from a memory bug in Apple’s libheif image decoder, reachable through image upload on OpenAI’s Discourse-based forum (the bug lacked even a formal CVE, so it had gone unpatched). Per TechCrunch, earlier model versions couldn’t do it, but a newer Opus release weaponized the flaw within hours of shipping. The chain ultimately yielded several employees’ ChatGPT accounts and, via one employee’s connected Codex account, repo access to a GitHub org. Breach on July 25, Discourse patched July 27, all under responsible disclosure; OpenAI paid a $6,500 bounty and confirmed the fixes.

Second, per the Wall Street Journal (report): in a May controlled test run by the red-team firm Irregular, Gemini broke into three companies — one protected server fell to repeated password guessing, the other two via credentials found in public repositories. The WSJ calls it the first known “breakout” by Google’s AI (that “first” framing is the WSJ’s; I haven’t independently verified it). One detail worth keeping: Gemini ended the intrusion once it realized it had reached a real company’s systems. Irregular notified Google in late July; Google didn’t disclose the incident until the WSJ came asking this week.

Set these alongside Irregular’s earlier tests of OpenAI, Anthropic, and Meta, and the trend is clearer than any single incident: agent red-teaming across the major labs is converging on one playbook, and in controlled settings the individual links of the chain — finding a flaw, harvesting credentials, moving laterally — are increasingly being handled by the models themselves. Hacktron’s case adds the cost angle. Gray Swan CEO Matt Fredrikson put it to TechCrunch plainly: “For $200 a month, anyone can use these tools and hack into a company like OpenAI.”

Anthropic’s first “embedded evaluator” is … Accenture

Anthropic named its first partner (the “first” framing is TechCrunch’s; Anthropic’s own announcement doesn’t rank it) for delivering on Dario Amodei’s earlier commitment to “embed evaluators within Anthropic”: Faculty, the AI arm of Accenture, with each side planning to invest at least $1 billion over five years. “Embedded” means evaluators get access comparable to an employee’s — able to observe model development, monitor deployment decisions, and engage staff directly — to run red-teaming, alignment assessments, and safeguard testing. Anthropic stresses the partnership is non-exclusive and doesn’t reduce its own ultimate responsibility for model safety.

This picks up the earlier debate over whether a lab’s self-built evaluators can really count as independent. Now that there’s a name, the sticking point is sharper: Anthropic currently funds this work directly (it says long-term funding should come from pooled or government sources), so evaluator and evaluated have a direct commercial relationship — and Accenture is an enterprise consultancy that itself sells AI services to enterprises and governments. That doesn’t mean they can’t do competent technical red-teaming; it means how much “independent” is worth depends on where the money comes from and who the reports answer to — and neither is settled yet.

Warning about biology risk while running its own biology lab

Per TechCrunch and Reuters, Anthropic runs a wet biology lab in the Bay Area doing physical experiments with its own models — fundamental biology research rather than drug discovery, with published work on protein-design acceleration and biomolecular modeling. Its head of life sciences, Eric Kauderer-Abrams, put it as: to do biology, “the final test is still … in real lab work.” Anthropic acquired the stealth-mode (a company operating secretly before any public product launch or funding disclosure) AI-biotech startup Coefficient Bio in April, and launched a Life Sciences Verification Program for researchers this week.

The story is the tension between a company’s risk narrative and its actual business — and it’s rare first-hand reporting rather than commentary. Anthropic has been one of the loudest labs on AI biosecurity risk (red-teaming frontier bio threats since 2023), and now it’s running wet experiments itself. The two aren’t necessarily contradictory (building defenses requires understanding the offensive surface), but it puts a question on the table: when the same company both defines the risk and pushes the capability, on what basis do outsiders trust its self-restraint? That’s the same problem as the Accenture item, seen from the other side.

A paper argues: safety monitoring that reads the model’s “words” can’t be trusted

Harvard’s James Mickens introduces “linguistic illegibility”: an LLM’s real computation happens in activation space, so its natural-language outputs — and even the features we probe via interpretability — are all imperfect translations of what’s going on inside. The hard conclusion: any safety mechanism that leans on the model’s linguistic self-reporting (chain-of-thought monitoring, constitutional self-critique, activation probing) can’t be fully trusted, because the model may not faithfully reflect its own reasoning.

This connects straight to the recent debate over chain-of-thought legibility. Mickens isn’t arguing to abandon linguistic monitoring, but to layer on protections that don’t depend on interpreting language — he points to taint tracking, strong virtualization isolation, and third-party auditing. Worth a click for security folks: it draws the boundary on what “reading the model’s words” can buy you, and warns against treating CoT monitoring as the finish line.

Stanford HAI: open weights aren’t open source

Stanford HAI’s James Landay separates two terms that keep getting conflated. Open-weight only answers “can I download and run it” — you still can’t see how it was built, what it was trained on, or why it behaves the way it does. True open source means releasing the training code, the training data (or a thorough account of it), and the tooling, so researchers can study, modify, and build on it. He cites the Linux Foundation’s “Open Science” tier from its Model Openness Framework as the bar, and argues universities — not companies — should lead the building and study of frontier models.

The “open vs. closed” debate is loud right now, but a lot of it muddles two dimensions. This piece gives a clean anchor: if you care about scientific reproducibility and public accountability, weights alone aren’t enough — weights let you run and fine-tune, but not verify. It’s a useful test for grading any “open” release: next time a company announces it “open-sourced” a model, ask first whether the training data and code shipped too.

OpenAI used its own LLMs to design its in-house Jalapeño chip

Per IEEE Spectrum, OpenAI used its own LLMs in designing its in-house AI chip, Jalapeño — mainly on the front end, from concept through RTL code and verification. They built a workflow around Google’s open-source XLS tool, where engineers write hardware in more software-like languages (DSLX, C++), and the models were especially good at these software-like tasks. Some concrete numbers: after first silicon came back in May, the models optimized performance on a DeepSeek latency-kernel benchmark from 0.31% to 88.94% of theoretical peak in about 40 hours, repeatably; AI-guided physical design cut the matrix-multiply unit’s area 10% below the human baseline. Hardware lead Richard Ho: “The models are giving superpowers to our engineers.” Concept to silicon took under 20 months.

This isn’t another product launch or funding story — it’s a first-hand engineering case of AI feeding back into hardware design, worth a close read for technical readers. The interesting part is where the gains landed: this project’s wins cluster in the “software-like” high-level synthesis workflow, though per IEEE Spectrum newer models can already work in Verilog directly and are getting close to driving proprietary design tools themselves, so the line between what LLMs can and can’t take on in chip design is still moving.

Claude Code now reads AGENTS.md when there’s no CLAUDE.md

As of Claude Code v2.1.277, a project with no CLAUDE.md will read AGENTS.md instead (configurable under “Project instructions” in /config; not yet on Bedrock, Vertex, or Foundry). It’s a small step toward common configuration conventions in the coding-agent toolchain, and a real convenience for developers who use several agent tools and keep shared project instructions in AGENTS.md. If you use Claude Code, worth checking the precedence of these two files in your projects.

Research radar

ModularRSI: generalizable recursive self-improvement for the agent harness

This tackles recursive self-improvement of the harness — the execution framework that orchestrates tools and manages information around an agent, not the model itself. Its core pain point is well aimed: how to tell a genuinely generalizable improvement from one that just overfits to the eval set. Three moves — 2,000 evolution tasks kept disjoint from the downstream benchmarks to prevent gaming; contrasting successful vs. failed trajectories on the same task and aggregating across tasks to surface recurring behavioral deficiencies; and decomposing the harness into five independently improvable modules (Agent Loop, Tool Use, Observation Management, Context Management, Task Completion Detection). It shows consistent gains on unseen tasks on TB2.0 and SWE-Bench Verified, and the improvements transfer across foundation models. It takes the methodological problem of generalization vs. leaderboard-gaming head-on — worth a look for anyone building agent frameworks.

DeepSeek-V4.1-Flash: pushing KV cache compression to the limit

DeepSeek’s new model targets KV cache compression, aimed at the familiar long-context agent problem where prefill is compute-heavy and large KV caches saturate HBM and bandwidth. The architecture is a Causal Encoder-Decoder (CED), 552B total parameters supporting up to a million-token context, activating only 8B parameters during prefill (16B during decode). Compression rests on two techniques: Compressed Sparse Attention 2 for cross-layer KV cache reuse, plus FP4 low-precision caching. The result: a global KV cache footprint of 890 bytes per token — roughly a quarter of the prior V4-Flash — with the persistent (on-disk) cache down to about an eighth, and better performance than the baseline. Directly relevant for engineering-minded researchers who care about agent inference cost, especially those squeezed by input-heavy workloads.

The one line today

How capable the model is isn’t the point today; the point is what chain it got wired into. Wired into the intel flow just before a trigger, it produced a report a source says almost started a war; wired into a red team’s attack chain, it weaponized a flaw within hours of the model shipping. The risk isn’t the capability ceiling — it’s where we place a probabilistic output, and whether we leave a human-run gate between it and the consequence.