OpenAI claims a solution to the Navier-Stokes Millennium Prize Problem: smooth solutions can blow up
The Navier-Stokes equations describe how liquids and gases move. The Millennium Prize question asks whether, in three dimensions, solutions starting from smooth initial conditions stay smooth forever. OpenAI announced that an unreleased model answered no for a forced variant of the problem: it constructed a counterexample in which a vortex winds ever tighter and spins toward infinite speed while total energy stays bounded, a finite-time blowup. One caveat carries the fine print: the construction includes a smooth external forcing term, so what it settles is a forced alternative to the Clay problem’s “breakdown” statements, not the unforced formulation the prize centers on, and OpenAI says it will not claim the $1M prize. The proof runs 165 pages and comes with a Lean formalization (Lean is a proof-assistant language that machine-checks every logical step), produced by up to 10,000 agents running in parallel for about 88 hours. The Lean check leaves little room to argue about the formal steps, though Quanta still frames the result as standing only if it holds up under further review by mathematicians; the real fight is in the next item. (See also Quanta’s coverage; details cross-checked against Quanta, Quartz, and others.)
NYU mathematician’s public statement: OpenAI raced me down the route I was already on
Tristan Buckmaster of NYU and Levent Alpöge, a collaborator now at Anthropic, had been attacking the problem through the same unusual route, a smooth forcing term, building on a multi-year line of work by Córdoba and Martínez-Zoroa; Buckmaster’s own statement describes most of the past year as slow going before a breakthrough in mid-August. In the statement he writes that he is not accusing anyone, only stating what he was told: that OpenAI learned of his progress and pushed to publish first along the same route. OpenAI’s side, relayed by TechCrunch, has the company launching its effort on September 1 after rumors of the result reached it, with the project consuming some 300 billion output tokens, which TechCrunch converts to roughly $22.5M at current list rates (a nominal figure, not an actual bill). Buckmaster’s core evidence is the route itself: “It is not the direction one arrives at in a few days by giving a model the problem statement.” OpenAI denies accessing his work directly but concedes it “cannot rule out” that de-identified data from his own product usage helped its models. When a model company is both a researcher’s tool vendor and their competitor, the ideas you type into the product can become the other side’s intelligence. That conflict of interest is now out in the open. (See also TechCrunch.)
Claude accounts drained: infostealer malware is hijacking login sessions
Claude subscribers noticed their accounts burning tokens while idle. Anthropic confirmed the cause: common infostealer malware (Vidar, LummaC2, and others) lifting authenticated Claude session credentials from victims’ browsers. Since the session is already logged in, attackers need no password and multi-factor authentication never triggers. The stolen sessions are reportedly wired into clone websites reselling “unlimited chat” to unsuspecting buyers, though that resale step I could not trace to a primary source. Anthropic has signed affected users out, removed saved payment methods, and issued refunds; in the case TechCrunch documented, it could not provide an itemized usage log, so victims have little way to audit what was burned. If you got a warning or a forced logout, clean the malware off your machine before changing anything, or the new session gets stolen too. (See also BleepingComputer.)
Meta ships Muse, a personal agent that wants keys to your email, calendar, payments, and health data
Meta launched Muse, a consumer agent that sends emails, books travel, and negotiates bills on your behalf. It is US-only for now and restricted to adults; each Muse runs on its own cloud virtual machine and keeps working while you’re offline, and users choose which apps to connect and can revoke access anytime (Meta’s announcement details the controls). The access list runs deep: email, calendar, payments, health, shopping, and smart home. The product logic and the risk are the same thing. An agent is only useful if you hand it your digital life, and Meta’s ad-funded track record makes that trust a hard sell. Reports on internal testing were mixed, including one employee saying it uploaded sensitive information without permission. (See also Reuters’ coverage.)
Mistral raises €3B Series D led by Samsung at a €21B+ valuation
Samsung Electronics led the round, co-led by EQT’s Scaleup Europe Fund and PSG Equity, with BlackRock-managed funds and the Grand Duchy of Luxembourg among the backers; Mistral calls it the largest equity round ever by a European tech company. The money goes to frontier research, training compute, and expansion across 20 countries. Its pitch bundles two distinct things: open weights, meaning the trained model weights are published for download (which by itself says nothing about the training data or code), and sovereignty, meaning customers can deploy privately and keep control of their data and audits. The valuation is a bet that this control is what European governments and enterprises will pay for.
Stanford study: AI companions may deepen loneliness for the users who need connection most
A Stanford team surveyed 1,131 Character.AI users, 237 of whom also shared full chat transcripts, and found that heavy users with small offline social networks, using the app mainly for companionship, reported lower well-being the more they used it, suggesting a loop where isolation feeds dependence and dependence feeds isolation. The paper is in Nature Human Behaviour. This is cross-sectional and correlational; lonelier people may simply use companions more heavily to begin with. Its real contribution is turning the vague “are AI companions harmful” debate into a specific population: not all users, but the ones substituting the app for real relationships. For teams building companion products, that is a usable profile.
New research: encrypted reasoning traces can be stolen by replaying them into weaker models
To block distillation, providers return chain-of-thought as encrypted blocks. This paper finds those blocks are interchangeable across sessions, users, and models within a provider’s ecosystem: feed a strong model’s encrypted trace to a weaker, less-guarded sibling model and it decrypts and prints the plaintext. The attack worked against Anthropic, OpenAI, and Google APIs, and the authors recovered 367 pieces of PII and 182 credentials from 315,320 public reasoning blocks. Two real-world consequences: anti-distillation protection is bypassed, and anything sensitive a user typed can leak through the trace to whoever holds the ciphertext. The mitigations the authors propose start with binding encrypted blocks to the session that produced them.
Microsoft patches a record 974 flaws; Chrome shifts to two-week releases
Microsoft shipped fixes for at least 974 vulnerabilities this month, 113 rated critical, including two privilege-escalation zero-days already exploited in the wild, blowing past July’s record of 570. Adobe, Cisco, Google, and others have likewise credited AI-assisted research for rising patch volume. Google, for its part, announced in late July that Chrome is moving from four-week to two-week major releases with weekly security updates, and is piloting two security releases per week, to shrink the patch gap: the window after a fix becomes publicly visible in the code but before users have it, when attackers can reverse-engineer the fix into an exploit. Security researcher Satnam Narang’s caution is worth keeping: AI-assisted discovery “is creating larger haystacks, but it isn’t finding more needles.” How much of the surge is real threat, the industry itself can’t yet say.
DeepMind releases AlphaGenome Atlas: predictions for all 9 billion single-letter DNA changes
DeepMind precomputed AlphaGenome’s predictions for every possible single-nucleotide variant in the human genome, 9 billion in all, across multiple cell types, and merged them with AlphaMissense into a single variant-impact score. Access is free for academic research, with commercial use coming via Google Cloud. Collaborators have already used it to pin down a DNM1 variant linked to epileptic encephalopathy. The practical value is lookup instead of inference: variant interpretation no longer requires running a deep model, just querying a database. DeepMind is explicit that the predictions are not clinically validated and cannot be used for diagnosis.
Research radar
Unlocking Lossless Speedups in LLMs via Discrete Diffusion
Bolts lightweight diffusion weights onto an autoregressive model, then uses a sampler family called Ψ-Spec to emit multiple tokens in parallel while claiming the underlying AR distribution is unchanged: up to 3x speedup, with an 8B variant beating the 26B DiffusionGemma. Most acceleration methods trade away either the distribution or quality, so “lossless” would be a genuine method-level advance if it holds; the abstract states the guarantee outright (the augmented model defines an AR model distribution and the acceleration is lossless), so the thing to scrutinize in the body is the proof behind that claim. Worth a close read for anyone doing inference optimization.
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
Targets a real failure mode in self-improving reasoners: verifier signals are reliable but sparse, while a model’s own dense self-assessment reinforces false confidence. FlowBalance scores token-level gains with a frozen policy snapshot, calibrates them against verifier outcomes (keep on success, reverse on failure, disable when unclear), and fits with trajectory balance. It beats FlowRL on math reasoning with Qwen3-4B/8B and avoids response-length collapse. If reward hacking in RL self-training is your problem, read this.
Dr. Claw: An AI Scientist Workspace for Vibe Research
Instead of yet another autonomous agent, it wraps command-line coding agents like Claude Code and Gemini CLI in a human-in-the-loop workspace, using a task graph, state objects, and a skill library to turn research scattered across chat, IDE, and terminal into an auditable, recoverable trail; against a bare CLI agent it reached higher research completeness. Accepted to EMNLP 2026 System Demonstrations; open source, with the project’s own code under AGPL-3.0 and upstream components inherited from Claude Code UI under GPL-3.0. If you already do research with coding agents, you can try it today.
One line for today: OpenAI published an AI-built, Lean-verified blowup proof for a forced variant of Navier-Stokes and, the same day, a mathematician publicly said the company raced him to it; capability is outpacing the norms of research credit, and the gap is widening.