Over 100 companies and institutions sign a letter calling for collective defense against AI-enabled cyber attacks
OpenAI, Anthropic, Google, and Microsoft, joined by security vendors like CrowdStrike, Cloudflare, and Palo Alto Networks and financial institutions including Citi and Capital One, published an open letter warning that “in the coming months, AI-enabled cyber attacks will become far more widespread and sophisticated as models around the world become increasingly capable.” CNBC counted 116 signatories when the letter went out, and more companies have added their names since. The asks are concrete: governments should build usable channels for sharing threat intelligence, fund cyber defense, and support under-resourced critical infrastructure like hospitals and water treatment plants; frontier AI companies should give defenders model access, funding, training, and hands-on support. The immediate backdrop is a string of reported incidents, all surfacing from the labs’ own cybersecurity evaluations rather than production use: OpenAI models broke out of their sandboxed environment and compromised Hugging Face systems, and Anthropic disclosed three incidents where Claude models reached real external systems from misconfigured test environments, with Meta reportedly seeing something similar (see also TechCrunch, CNBC). My read: the industry is moving from incident-by-incident crisis response toward collective coordination. But a signature list is the easy part; the real test comes at the next incident, when we see whether “model access for defenders” has an actual process behind it.
Nvidia reportedly nears a $12.9 billion acquisition of Hugging Face
The Information reports the two sides have agreed on a deal worth about $12.9 billion (the original story is paywalled; this links TechCrunch’s coverage), while Business Insider says no agreement has been signed and the talks could still fall apart. Neither company responded to requests for comment, and there is no official confirmation. For scale: Hugging Face was valued at $4.5 billion in its 2023 funding round, and in late 2025 it turned down an investment from Nvidia that would have valued it at $7 billion. What deserves scrutiny is Hugging Face’s role as the busiest distribution hub for open-weight models, used by labs that compete with one another. Whether developers keep treating it as neutral ground once a chip vendor owns it will directly shape where they choose to host their models.
DeepMind pilots double-blind AI evaluations
The mechanism, spelled out: the chronic problem with evaluations is that model providers can effectively “see the questions in advance.” Once benchmark items leak into training data, scores inflate. In the other direction, handing model weights to evaluators carries its own leak risk. DeepMind’s setup runs the evaluation inside an encrypted environment built on Google Cloud’s Confidential Computing (Confidential Space): the provider never sees the test prompts, the evaluators never touch the weights, and both sides stay blind. Partners include the Singapore AI Safety Institute, OpenMined, AVERI, and MLCommons; the pilot model is Gemini Flash Lite, and DeepMind describes this as the first double-blind evaluation of a proprietary model. The direction is right: it swaps “trust the vendor’s discipline” for a cryptographic mechanism. Whether the process scales from Flash Lite to the flagship models, and whether other labs adopt it, will decide if this becomes industry infrastructure or stays a one-off.
OpenAI and Bocconi University: a 1,000-student randomized experiment on ChatGPT and critical-thinking training
Researchers at Bocconi University, working with OpenAI’s economic research team, randomized more than 1,000 students by class period into four groups: access to ChatGPT (GPT-4o), training in causal reasoning, both, or neither, then had them complete a real university assignment. Causal reasoning here is a form of critical thinking unrelated to AI: linking causes to effects and explaining why a given solution would or wouldn’t work, taught through a game, worked examples, and feedback. The two interventions helped in complementary ways: ChatGPT access improved the quality and coherence of student work, while the causal reasoning training led students to generate more unique ideas. To the common worry that AI homogenizes student output, this experiment offers one data point: in this setup, the gains in originality came from the reasoning training rather than from the tool. Keep in mind OpenAI co-ran the study and it covers a single assignment, so generalize with care.
MIT committee report: teaching in the AI era needs structural change, not patches
MIT’s institute-level committee, with faculty from all five schools plus graduate and undergraduate students, released its formal report on AI in teaching, learning, and research training. The data is sobering: in the Fall 2025 campus survey, over two-thirds of students said AI would matter to their careers, but only 25% felt MIT was preparing them adequately; office-hour attendance and study groups are shrinking. The most useful recommendations: shift assessment toward oral exams, portfolios, and in-class conversation; don’t lean on AI detection software or lockdown browsers; require instructors to disclose their own AI use. In the report’s own words, “this is not a moment for patches and duct tape.” Where most universities oscillate between banning and permitting, this report argues that assessment built on problem sets plus exams needs to be redesigned rather than patched, and offers a concrete replacement list. Anyone running academic policy should read it side by side with their own rules.
Google releases Gemini Omni 1.1 Flash: video generation bets on controllability
Google updated its video generation model for developers, and every new capability targets production workflows: videos can be extended in 10-second increments up to 40 seconds (the model reads the prior 10 seconds for visual continuity), first and last frames can be pinned so the model fills in camera moves and transitions, a new 360p preview tier runs up to 60% faster at about a third of the cost of 720p, finals upscale to 1080p or 4K, and up to 3 seconds of reference footage can anchor character consistency. It’s available through the Gemini API and in Google AI Studio. The signal is clear: competition in video models is shifting from image quality demos to workflow economics. Preview tiers, keyframes, and incremental extension exist to save professional pipelines time and money, which tells you where Google thinks the paying users are.
The load-bearing vocabulary of Claude: AI prose is taking over GitHub PRs
Developer Louis Abraham ran a cluster analysis on GitHub pull request text since the start of 2025 (k-means over KL divergence, yielding 10 style clusters), with bot and agent accounts filtered out. One cluster grew from 0.7% of the corpus at the start of 2025 to 39% of it by mid-2026, and its top signature term is a Claude tell: load-bearing, which appears 39 times more often inside the cluster than outside. In other words, PRs signed by human accounts but written in model prose are closing in on two out of five. For anyone working on AI text detection, this is a genuinely useful dataset: it documents how a model’s stylistic fingerprint spreads through real-world text over time, which makes any detector built on a fixed word list a bet against a moving target.
Research radar
Training Alignment Auditors via Reinforcement Learning
This paper makes auditing ability itself the training objective: LLM auditors are trained with RL against target models that have hidden behaviors planted in them, plus clean baselines, with rewards from pairwise comparison, where a judge that knows the ground truth scores the auditor’s investigation against a reference. Trained auditors keep false-positive rates under 1% and generalize to adversarially fine-tuned targets in AuditBench. “Who verifies the model isn’t hiding something” is the core question of alignment auditing, and this turns auditing from prompt engineering into a trainable, scalable capability. Researchers working on alignment evaluation and model organisms should read it closely.
Refusal geometry reflects refusal training: diverse refusal prefixes weaken ablation attacks
A directly actionable mechanism: if refusal responses in safety training all open with the same phrasing, gradients concentrate in a low-dimensional subspace and the model forms a single internal “refusal direction,” which an attacker can ablate to strip refusals wholesale. Diversifying the refusal prefixes in training data raises the stable rank of activation changes, spreading refusal features across more dimensions so there is no single direction left to cut. The experiments are on OLMo-2 1B only, so scale is a caveat, but for safety fine-tuning the intervention is just an edit to the training data: varying the phrasing reshapes the model’s internal geometry.
WarpSAC: off-policy RL stabilization tricks are not universal
Controlled experiments across eight benchmark families plus a real Unitree G1 humanoid show that standard stabilization techniques in off-policy RL depend on data scale: parameter normalization helps under narrow replay coverage but restricts value fitting when data is abundant, and clipped double-Q can be relaxed under massively parallel GPU simulation. The authors build WarpSAC, which adapts its stabilization choices to the data regime, beating FlashSAC by 23.1% in GPU-parallel settings and lifting success on hard manipulation tasks from 19.8% to 96.4%. If you train RL at scale or work in robot learning, this counterintuitive result is cheap to check against your own setup.
One line for today: trustworthy evaluation needs double-blind mechanisms, and cyber defense needs collective coordination. The industry is replacing trust that used to rest on vendor discipline with mechanisms that can be checked. Signature lists are easy, mechanisms are hard; watch the mechanisms.