OpenAI says an internal model solved 100+ open math problems. Verifying them is not its new advisory group’s job

On September 21 OpenAI announced a Mathematics and AI advisory group of nine mathematicians, among them Fields Medalists Timothy Gowers and Martin Hairer, and Edward Witten, a physicist who also holds a Fields Medal. In the same announcement OpenAI claimed its internal model has resolved more than 100 long-standing open problems across most areas of mathematics, on top of the Navier–Stokes result it published earlier in September, which is itself still awaiting independent acceptance. Two things sit inside that claim. First, the “100+” is OpenAI’s own count; the announcement points to no peer review and never defines what counts as an “open problem” or who checked the answers. The group’s own site (agmai.org) states no number at all, saying only that OpenAI “reported” a large number of significant results. Second, the group’s mandate covers assessing how significant such results are and advising how and when to publish them — significance assessment, not peer review; neither OpenAI’s announcement nor the group’s site says who, if anyone, verifies the results themselves. The group’s site stresses that it operates independently of any AI company and that members take no payment, and OpenAI’s announcement adds that the group will not advise it on the pace of its internal research. Context sharpens it: on September 11, 25 Fields Medalists signed an open letter warning that labs racing to crack famous problems crowds out mathematical work, and one of the nine group members, Martin Hairer, is among its initial signatories. My read: when AI claims a scientific result, the load-bearing question is who checks it and how, and this arrangement hands “how important” to experts while leaving “is it correct” with the company.

Amazon blocked Meta’s shopping agent Muse

When users pointed Meta’s personal AI agent Muse at Amazon to place an order, Amazon returned an error: “Continued access by an unauthorized AI agent violates Amazon’s Conditions of Use, to which our customers have agreed.” Amazon is invoking its Conditions of Use here; the report does not say how the block is implemented technically. TechCrunch’s reading — its own, not a stated rationale from Amazon — gives two reasons: Amazon runs its own competing AI and owes a rival’s agent nothing on its storefront, and it worries about the failed orders and customer complaints that even a low-hallucination agent will produce. Worth pairing with Muse’s momentum: per estimates from market-intelligence firm Apptopia (reported by TechCrunch), Muse’s first 12 days beat ChatGPT’s early mobile numbers on US-plus-Canada iOS downloads (1.8M vs 1.3M) and US daily actives (642K vs 231K). Those are third-party estimates, not Meta’s internal numbers. The real battleground for agentic commerce (AI placing orders on your behalf) is not model skill but platform access: whoever owns the shelf decides which agents get to transact.

Hackers reverse-engineered a Flock plate-reader camera and found it also profiles people and bikes

Anonymous hackers physically pulled a Flock camera off its pole, copied nearly all of its stored data, and handed the files to 404 Media (reporters Joseph Cox and Dhruv Mehrotra, September 16) and WIRED. Flock sells automated license-plate reading (ALPR), but the software they recovered classifies far more than plates: people, vehicles, plates, and bicycles. The logs held over a million images across several weeks, and the vision software sometimes isolated identifying details such as bumper stickers and an American-flag patch on a motorcycle saddlebag. They also found a security hole: an unencrypted partition stored the key that unlocks a protected one. (The 404 Media piece is paywalled; see the archive. Schneier’s blog summarized it.) What a surveillance device advertises and what its firmware can actually recognize can be a wide gap, and it took physically opening one to prove it.

Nathan Lambert: on open-weight models, Chinese labs really do lead

Lambert expanded his US Congressional testimony into an essay with a blunt conclusion: on open-weight models (weights that anyone can download), Chinese labs passed US open models roughly 18 months ago and now lead on benchmarks, downloads, adoption, and citations. He lays out the gap this way: Chinese open models trail the US closed frontier (OpenAI, Anthropic) by about 2 to 5 months, while US open models trail by about 6 to 9 months, which puts the US open-weight ecosystem behind both its own closed frontier and Chinese open models. He does not dodge the “China just distills US closed models” rebuttal, but he sizes it small: distillation explains maybe 1 to 2 months of the gap, and the rest he credits to faster release cycles, sharper task focus, and deliberate ecosystem strategy such as Alibaba investing to get Qwen cited in research. I’ve noted before that “open weights ≈ Chinese labs” is a recent impression, since Meta’s Llama was the first open-weight LLM to see mass adoption and the framing only formed in the year or two after DeepSeek. Lambert’s addition: inside that short window the Chinese lead is structural rather than copied, and what Llama has left is longevity, not leadership.

OpenAI calls for shared standards for AI’s “next phase”

The same day, OpenAI published a policy piece arguing that as AI takes on more autonomous research and even recursive self-improvement (AI improving AI R&D itself, potentially into a feedback loop), the field needs interoperable technical standards; OpenAI goes so far as to say such standards may be “as important to pacing the frontier as alignment research itself.” It names three concrete standard areas: evaluation, meaning how to measure how much autonomous AI research is happening inside a company; human-oversight triggers, meaning which automated research processes should force immediate human review; and incident classification and reporting. It names the bodies it wants to carry this (the US CAISI, ISO, the Frontier Model Forum, national AI safety institutes) and draws one clear line: these standards should not become licenses or mandatory approvals. It is a directional proposal, not drafted standard text, and how far it goes depends on who actually writes the clauses.

Tim Dettmers: frontier models on hardware you already own

Quantization researcher Tim Dettmers and his DLab ran an “Open Source Week” aimed at letting compute-poor academic labs run frontier-class models and autonomous research agents on hardware they already have, rather than on corporate infrastructure. Three pieces: an inference-serving framework with custom kernels, an agent harness for autonomous research, and a compaction technique for very long agent sessions (he calls it CliffCompaction). What makes large models fit on consumer devices is aggressive quantization, down to about 1.5 bits per weight, roughly a 90% memory cut versus 16-bit, enough to run tens-to-hundreds-of-billions-parameter models on a 24GB desktop GPU or a 128GB MacBook (the post names specific models and throughput figures, which I won’t relay one by one). The point of this line of work: frontier capability is sliding from “only a data center can run it” toward “a person or a small lab can run it locally,” and running it on your own hardware means the data need not leave it.

Linear: AI writes code so fast it turned CI into the bottleneck

Linear’s engineering team wrote up how AI coding agents made writing and committing code explode while validation lagged, turning CI (continuous integration, the pipeline that runs tests and checks on every commit) into the new bottleneck; their test suite nearly quadrupled since January, dragging wait times and machine cost up together. The fixes are specific: move to third-party runners (about 34% faster jobs), swap the TypeScript compiler for tsgo (typechecking down 73%), adopt Oxlint, use sparse checkouts to cut change-detection from 26s to 8s, move cache writes off the merge path (42s saved per PR), pre-install dependencies in CI base images (pnpm install from 44–73s down to 16–18s), consolidate 7 checks into 2 jobs (about 87,000 runner-minutes a month), and shard tests from 4 to 8. PR wait dropped from over 6 minutes to just over 5. A recurring pattern fits again: once AI makes writing code cheap, the cost does not vanish, it flows downstream to validation and review.

Foremerge: catch intent conflicts between parallel agents before they write code

This open-source tool (a v0.5 MVP written in Rust) targets a hazard of running several coding agents in parallel: two agents can make edits that never overlap textually, so Git merges both cleanly, yet their intents collide. The canonical case: agent A replaces PaymentService wholesale with StripePaymentService while agent B is adding PayPal to PaymentService, and after the merge B’s work is orphaned. Foremerge has each agent declare intent up front (which symbol, replace vs extend, a one-line summary) into a shared SQLite store, then applies deterministic rules to compare operations on the same target and flags conflicts. It only advises, never locks files, has no published benchmarks yet, and its detection is heuristic. Multi-agent coordination failures are a thread I keep following, and this is a concrete take on moving conflict detection to before the code is written.

Jev: a model that emits probability-weighted decisions instead of text

TypeSafe AI released Jev, which it calls the first “System One Model” (the name nods to Kahneman’s fast “System 1”). Unlike a normal LLM it does not generate text; it outputs typed, calibrated decisions: give it a described situation plus a set of allowed actions, and it returns a decision bound to a probability and a confidence. Two mechanisms stand out: it is non-autoregressive (no token-by-token generation; it samples the whole output in parallel), and it replaces RLHF with RLCD, reinforcement learning for calibrated decisions, which TypeSafe says trains the stated confidence to track the actual hit rate. Because outputs are constrained to predefined types, it claims a zero structured-output error rate. TypeSafe’s self-reported numbers are 70–500ms responses at about $0.0004 per case, which it advertises as 40 to 200 times faster and several hundred times cheaper than frontier LLMs; all of that is vendor benchmarks, not independently verified. The thing to watch is the category itself: a model whose output is a decision value, not prose.

Spymarks: a sharper name for the kind of watermark that can identify you

This piece goes after the vocabulary. It splits “watermark” in two: visible, inspectable, removable marks that assert ownership stay “watermarks,” while invisible marks applied without your consent, surviving compression and re-encoding, that make a work traceable back to you get a new name, “spymark.” It gives mechanisms: images use frequency-domain pixel changes (it cites Google SynthID-Image embedding a 136-bit payload in a 512×512 image — room enough, the essay argues, for a 64-bit identifier that could link to user records); audio uses inaudible waveform modulation (the open-source audiowmark hides a 128-bit payload, with the embedding keyed by AES); text uses word-choice bias. This is the same worry I raised last month in my watermark piece: how finely the key is handed out decides whether a watermark is a provenance tool or a user fingerprint. What this adds is putting that worry into the name itself, with “spy-” stating the surveillance up front.

Research radar

IntBMoE: unbundling “participation, compute, storage” in mixture-of-experts

In a standard MoE (mixture-of-experts, where a model keeps a set of “expert” subnetworks and routes each token to only a few), three things move together: how many experts contribute knowledge to the output (participation), how many are actually computed (compute), and how many expert-sized parameter sets must be built and stored (storage). This paper pulls them apart so each can be set on its own: a lightweight hypernetwork merges all expert bases into one composed expert for full participation, a router still sends each token to only a few blocks for sparse compute, and a codebook caps the parameters that must be materialized for bounded storage. The result is a 2.4% lift in a click metric (UVCTR) in a live A/B test in AMap’s recommender under a 60ms latency budget, plus consistent gains over sparse and dense MoE baselines on vision tasks. Worth a look if you build large-scale recommendation or low-latency serving, or work on efficient architectures.

Verifiable Social Reasoning: manufacturing ground truth for a task that has none

Social reasoning is hard to evaluate because it is subjective and lacks ground truth. The Fuse framework sidesteps that with a multi-agent simulation: a target agent with a hidden motive interacts with other agents, including one standing in for the user, who then consults the assistant under test. Because the hidden motive is set by construction in the simulation, the correct answer is known in advance, so the assistant’s judgment can be scored against that planted motive instead of against subjective human ratings. The team evaluated 12 LLMs over about 21,000 examples, with four findings: user mediation compounds difficulty, models are systematically sensitive to biased framing, models need more detail than humans to guess right, and longer conversations do not reliably improve accuracy. Useful for anyone building social-advice or consultation assistants, or studying framing effects and intention inference in alignment work.

One line for today: When AI starts claiming it solved hard math or made a scientific discovery, the question that carries weight is not how many it solved but who checks whether it’s right. OpenAI’s advisory group rates how important the results are; who verifies that they hold, the announcement never says.