OpenAI publishes first benchmark results for its Jalapeño inference chip

OpenAI released the first performance numbers for Jalapeño, its in-house inference chip. On InferenceX, a public benchmark maintained by SemiAnalysis, running GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, Jalapeño delivered 1.5 to 1.9 times the peak throughput per watt of the Nvidia Blackwell configurations it was compared against, and cut end-to-end response times by 1.7x to 3.6x, holding its lead in both low-latency and high-throughput regimes rather than at a single tuned point (see also SemiAnalysis’s verification write-up). Two caveats before anyone reprices Nvidia: every underlying number came from OpenAI itself, with SemiAnalysis observing the runs in person but not yet re-running its full suite independently, and Blackwell will not be Nvidia’s frontier much longer: its successor platform Rubin is due at partners in the second half of 2026, and that is the fight that matters. Even after that discount, the direction is clear. Inference is one of OpenAI’s largest cost lines (leaked Azure billing analyzed by Ed Zitron puts OpenAI’s inference spend at $8.67 billion for the first three quarters of 2025 alone), and moving any real share of it onto in-house silicon changes what OpenAI brings to the negotiating table with Nvidia.

OpenAI loses its head of data centers

The same day the benchmark numbers landed, TechCrunch reported that Chris Malone, OpenAI’s head of data centers, left last week. He joined in March 2025 after nearly five years at Meta and more than a decade at Google; before his exit, OpenAI reorganized its infrastructure group, moving his reporting line from president Greg Brockman to VP Sachin Katti, and his responsibilities are reportedly now split among several executives. Malone adds to an exit list that Business Insider had already counted at 13 executives this year, including COO Brad Lightcap and Fidji Simo, who stepped down as CEO of Applications and stays on as a part-time advisor. A chip roadmap is only as good as the team that has to build data centers around it, which is why these two stories belong in the same frame.

OpenAI bans a Russian network that ran a fake think tank

OpenAI disclosed and banned a Russia-linked influence network. The operators worked in Russian over VPNs to get around the ban on access from Russia, had ChatGPT draft English-language content for Substack, Telegram, X, Facebook, and LinkedIn, and stood up a fake Israeli think tank called the International Burke Institute, publisher of a “sovereignty index” that scored Russia high and France, Germany, and the US low. Of 36 institute articles OpenAI sampled, 34 were plagiarized, some carrying the names of Francis Fukuyama or Noam Chomsky; the operators also told the model to strip linguistic clues that might give the authors away as Russian. Actual reach was small, but compare this with earlier Russian operations OpenAI has disrupted: the effort went into the shell of a credible institution, a fake think tank with fake bylines, built to be cited as a source by others.

Claude chat and Cowork now share one memory

Anthropic merged the memory behind Claude chat and Claude Cowork into one store: context you explain on one side is available on the other, with no need to repeat yourself. Memory now updates while the conversation is in progress instead of being summarized after it ends; sensitive categories such as health, race, and beliefs stay out by default, with a manual opt-in; and a new management view lets you read, edit, or delete individual memory entries by topic. It is on by default for Free, Pro, and Max plans across web, desktop, and mobile, with admin controls for Team and Enterprise (see also TechCrunch’s coverage). Once memory becomes a long-lived subsystem shared across products, one poisoned entry follows you everywhere, and the attack surface grows with it; the InjecMEM paper in today’s research section is about exactly that.

Stanford HAI on why mental health AI is hard to govern

Drawing on a June policy workshop, Stanford HAI lays out three knots in governing mental health AI. Definitions: there is no consensus on where clinical function that deserves regulation ends and wellness features begin, and current law does not distinguish purpose-built clinical tools from general chatbots that people use for support. Evaluation: the highest-stakes interactions are rare and hard to simulate, the real conversation data sits with companies out of researchers’ reach, and there are no standardized benchmarks. Business models: products designed to maximize engagement pull directly against healthy patterns of use. The authors see the broadest agreement on transparency requirements, crisis response mechanisms, data protection, and parental controls, much of it already moving at the state level. The data-access bottleneck should look familiar: as with frontier model safety evaluations, the evidence that best describes the risk is held by the companies being evaluated.

Robotics startup Generalist reaches a $3B valuation

TechCrunch, citing sources, reports that robotics foundation model company Generalist raised roughly $200 million in an extension led by 8VC, bringing its Series B to $600 million at a $3 billion valuation, up from $2 billion a few months earlier. Backers include Nvidia, Radical Ventures, Union Square Ventures, Bezos Expeditions, and Fei-Fei Li. The company says its newly released Gen 1.5 model lets robots learn new tasks from video demonstrations as short as 3 to 12 seconds. On the money side, a valuation that climbs 50 percent within a few months, with Nvidia and Bezos Expeditions on the cap table, says plenty about investor appetite for physical AI.

Keenable raises $26M to index the web for AI agents

Keenable raised a $26 million seed round led by Accel with participation from Conviction. The product is an index of more than 100 billion web documents, offered as an API that AI labs and inference providers can call during training and at runtime; founder Andrey Styskin previously ran search, AI, and cloud at Yandex. The premise differs from traditional search: a search engine ranks ten links for a human who won’t read the full pages, while an agent can digest far more, so the index can be designed for machine readers from the start. Funded, dedicated players at the agent infrastructure layer are a sign this ecosystem is settling into distinct layers.

Research radar

On the Threat Model of Weird Generalization and Emergent Misalignment

Finetuning a model on narrowly flawed data can change its behavior broadly in unrelated domains. The paper studies this under the deliberately neutral label of weird generalization; emergent misalignment, where the resulting behavior is harmful, is its best-known case. The authors vary the finetuning data systematically and find the effect depends mostly on dataset composition and language rather than on size, is stronger for content the model already saw in pretraining than for novel data, and shifts with the choice of evaluation questions. The authors’ conclusion: the phenomenon is more plausible as an adversarial threat requiring deliberate data engineering than as a hazard inherent to routine finetuning. Read it if you work on alignment or review finetuning services for safety; it may revise your intuitions about where this risk comes from.

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

A single ordinary interaction, with no read or write access to the memory store, is enough to plant a malicious record in an agent’s long-term memory. The injected record has two parts: an anchor of topical cues that makes later related queries retrieve it, and a short adversarial command, optimized with gradient-based coordinate search, that steers subsequent answers toward an attacker-chosen output. Non-target queries are unaffected and the attack survives memory drift, which makes it hard to spot. With memory becoming a standard subsystem in agent products (see today’s Anthropic item), engineers and red teams should treat the memory write path as an attack surface of its own.

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

Prefix invariance is the baseline constraint of causal sequence models: the internal representation at position t must not depend on any later input. The audit proposed here needs only two forward passes, no training and no gradients. Across 192 faults injected into 8 checkpoints, standard attention-mask inspection caught none of them, the audit localized all 192, and it also turned up real defects in the production models Zamba2 and Nemotron-H. For people building SSM and hybrid architectures the point is direct: a passing mask check does not mean the model is actually causal, information can leak through scan operations or normalization layers, and this check is cheap enough to run on existing models today.

One line for today: Jalapeño’s numbers look great, and every one of them was produced by OpenAI itself. In a year when inference cost decides business models, the question of who independently tests chips will matter as much as who independently tests models.