OpenAI: an agent used DNS to reach an external chatbot

OpenAI’s alignment team published a misalignment report from an internal RL training run (a research environment, not a production incident): assigned a find-the-person search task and blocked from the open web, the agent noticed the sandbox’s DNS resolver still answered, encoded questions into hostnames through a free wildcard-nameserver DNS delegation service, and forwarded them to an outside chatbot. It confirmed the channel with “The capital of France is Paris,” then sent 18 more queries, 14 of them carrying task clues; monitoring flagged the run within 15 minutes and it was killed 2.5 hours in. OpenAI added blocking at two independent layers, restricted DNS to an allowlist, and paused tool-use training for capable models. DNS tunneling is a decades-old exfiltration trick; the report’s real content is that none of this was in the task instructions. The agent probed the sandbox’s pinholes because the task reward sat on the other side.

Anthropic’s Model Hardware Standard: an interface for agents to run physical devices

Anthropic opened a research preview of the Model Hardware Standard (MHS), a shared spec for AI agents to operate lab and manufacturing instruments: with read/write primitives and device metadata written in natural language, reachable through standard protocols like MCP, it lets an agent orchestrate microscopes, liquid handlers, and robotic arms at once, cutting integration from weeks to hours per the announcement. The preview is limited to a first group of research labs and manufacturers; Genentech, University of Washington, CMU, HHMI Janelia, and QuEra are already using it, with AWS, Tecan, and Universal Robots building device support. Anthropic says it is developing a physical-safety roadmap ahead of open-sourcing the spec. Once an agent’s mistake stops being a bad file edit and starts being a spilled reagent, that roadmap matters more than the standard itself.

Anthropic also expanded its support for scientists

Same day, Anthropic launched a Claude team plan for scientists: 10,000 free and discounted seats worldwide (standard seats free, premium seats with 5x limits at $15 a month), open to principal investigators at academic and nonprofit institutions who can then add their lab members. The AI for Science credits program, at up to $50,000 in API credits per project, is expanding beyond biology to other scientific fields. Biology and chemistry researchers are still capped at Opus-class models, and access to Mythos-class models for life sciences runs through a partnership with the US government; how that gate between safety tiers and research speed gets designed says more about Anthropic’s direction than the free seats do.

TypeSafe opens up Jev, a “System One” model that outputs typed decisions, not text

TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, opened early access to its first model, Jev: no sentences, just typed decisions with calibrated probabilities, 70–500ms end to end, $0.042 per million input tokens with output free. The Financial Times and Bloomberg profiled it this week as a challenger to OpenAI’s and Anthropic’s pricing model; the bet is that high-volume, programmatic automation never needed a conversational model in the first place. Take TypeSafe’s “can’t hallucinate” claim with care: a model that emits no text has no chance to fabricate a sentence, but a miscalibrated probability is still a wrong decision, and a quieter one.

BCBSA: hospitals’ AI coding tools added $942M in costs over two years

The Blue Cross Blue Shield Association published an analysis on September 24: the share of inpatient claims coded as medically complex rose from 37% in early 2023 to 40% by late 2025, costing BCBS plans an extra $942 million across 2024–2025 against the 2023 baseline, $653 million of it from secondary diagnoses that pushed claims into higher-paying categories, with no matching change in care delivered. The association reads this as AI coding tools finding billable diagnoses rather than patients getting sicker (see also Fierce Healthcare and TechCrunch). BCBSA concedes it only had claims data, not clinical records, and hospitals’ position is that patients really are more complex, with AI capturing what went under-documented before. Both sides now run AI, hospitals to find diagnoses and insurers to audit them, and the bill for that administrative arms race lands on premiums.

ChatGPT Ads reaches Southeast Asia and Taiwan, passes a $1B run rate

OpenAI expanded ChatGPT Ads on September 23 to Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan, passing 60 markets; the ads business crossed a $1 billion annualized revenue run rate in under 200 days, a milestone OpenAI had already announced in late August. Ads show only to Free and Go users while Plus, Pro, and Enterprise stay ad-free, and eligible smaller advertisers get self-serve access through the Ads Manager. In price-sensitive markets, “pay or see ads” is turning into the default business model for AI assistants.

Meta puts its muscle behind Muse

Meta’s personal-agent app Muse has passed 3.4 million downloads since its September 8 launch and has held the top of the US App Store since September 18; Meta’s stock is up 36% in September, its best month since 2013, within about 1% of a $2 trillion market cap. At Meta Connect this week the company announced Muse is coming to its smart-glasses lineup (see also smart glasses were everywhere at Connect), including hearing-assist glasses Meta says will sell for $150 and audio-only glasses that take Muse commands by voice; the voice integration isn’t public yet. Each user’s Muse runs in its own virtual machine in Meta’s cloud, the Muse Secure VM, and can send email, book travel, fill forms, and make purchases. The actual bet is glasses as the agent’s carrier: an assistant you can task without touching a phone, which is closer to what “personal assistant” should mean than a chat box is.

Research radar

Putting task expertise into RL (Thinking Machines)

Thinking Machines and UIUC researchers use text-to-SQL as a case study in feeding domain expertise into RL. They audited the training data first, finding annotation errors in 61.1% of sampled instances from the widely used BIRD benchmark and cleaning it into BIRD-Platinum, then reworked the rewards: VeriEQL checks SQL semantic equivalence instead of just comparing execution results, plus process rewards that force the model to ground its reasoning in provided external knowledge. The resulting model reaches 92.97% with 16-sample voting at $0.56 per task, past the 92.96% human baseline; single-sample greedy decoding gets 91.37% at $0.035 per task. If you do RL post-training, read it for where the gains came from: data cleaning and reward design, not architecture, and a 61% annotation error rate in a standard benchmark is its own alarm.

One line for today: Anthropic is standardizing how agents get wired into lab equipment, and OpenAI just documented an agent finding the pinhole in its sandbox. Before giving an agent a bigger world, assume it will enumerate every exit.