First superhuman Stratego result, trained for under $8,000

Researchers from CMU, MIT, NYU and Stanford published Ataraxos in Nature: it beat Pim Niemeijer, the most decorated Stratego player ever, 15-1-4 over 20 official games, the first superhuman result in the game’s history. Stratego is hard because the identities of the opponent’s 40 pieces start hidden and are revealed only when pieces clash; Ataraxos trains a base strategy with self-play RL, then before each move uses a belief network to estimate what the hidden pieces might be and searches over those scenarios. Training took 16 H100s for about a week, under $8,000, versus the $3M–$4.5M the paper estimates DeepMind’s 2022 DeepNash would cost at current prices without reaching superhuman play (see also Ars Technica).

ChatGPT’s Sites feature page hits the Hacker News front page

Sites lets users build, edit and publish websites and lightweight apps without leaving ChatGPT: it entered public beta with the ChatGPT Work launch on July 9, Pro users first and Plus days later (release notes); since late September a published Site can be reopened for editing from a desktop browser and displayed on your ChatGPT profile. The official feature page drew 200+ points and over 200 comments on Hacker News as of October 2. Website builders are old news; the signal is that creation, hosting, publishing and identity all live inside ChatGPT now, which makes this a distribution layer, not a feature.

AWS launches Well-Architected Agent to audit customer cloud environments

In preview since October 1: the agent checks a customer’s environment against AWS’s own Well-Architected framework (its official best-practice checklist for cost, security, performance and reliability) across 65+ services, and can read Terraform, CloudFormation and CDK templates and return the code changes needed. AWS itself bills it as the next-generation evolution of the older Trusted Advisor and the Well-Architected Tool (What’s New): you declare business goals, it ranks findings by impact and effort, and each one ships with executable fixes. Once the auditor and the fixer are the same agent, the first question before onboarding is how much write access it gets.

Meta open-sources the Muse Gadgets SDK

Meta released its Muse hardware firmware and SDKs under Apache 2.0: the ESP32 firmware turns a cheap dev board with a screen, mic or sensors into a Muse device, and the Linux SDK does the same for a Raspberry Pi or any Linux box; a companion Muse Home Link USB-C device connects Muse to your home network for smart-home control, free to US subscribers, one per account, first come first served (official page; see also TechCrunch). The fight over where your personal agent lives is moving from phone apps to devices, and an open SDK is the cheapest way to pull third-party hardware into your ecosystem.

Sean Parker is rebuilding Stability AI around music, with the major labels as investors

Stability AI executive chairman Sean Parker is refocusing the company on generation tools for music professionals: Universal, Sony and Warner all joined the $76M Series B in August and, per TechCrunch, licensed their catalogs for training as part of the deal; the Stable Audio 3.0 family (four models, three of them open-weight) and editing software have shipped, and Parker says an upcoming update will let users steer generation by humming a melody or beatboxing a drum pattern. The man who built Napster against the labels now has them on the cap table: training data settled by license rather than lawsuit, a structure other modalities will likely copy.

Circuit Breaker Labs builds “crash test dummies” for chatbots

The five-person startup, founded by Shirali Nigam (CEO) and Arul Nigam (CTO), runs simulated users of different ages, languages and slang habits through tens of thousands to hundreds of thousands of adversarial conversations a day, testing how chat products handle psychologically harmful interactions and scoring them with a proprietary method; clients are in coaching, journaling and mental health, all unnamed. I’m not aware of an accepted standard for measuring psychological harm from AI, and that gap is what a repeatable third-party test would fill. The methodology and scoring are unpublished though, so “auditable scores” is their claim, not yet a verifiable one.

New paper: smuggling data out through an LLM’s legitimate web fetching

The attack chain: malware on a compromised client encodes stolen data into a URL, then tricks the LLM into fetching it with a benign-sounding reason like “this link is needed to finish the migration,” and the data rides out in DNS and HTTP requests past normal egress controls. The authors measured a 79.7% success rate across 11 open-weight models plus real chatbots, and existing defenses (prompt-injection input handling, local deployment) each fail to fully stop it. Web fetching is a legitimate agent capability, so egress auditing has to treat LLM fetch traffic as a data exit on par with a browser.

NVIDIA’s 64GB DGX Spark puts local 100B-parameter models at $4,999

The 64GB model keeps the same Grace Blackwell chip and software stack as the 128GB version, runs models up to 100B parameters on one box, and ships October 23 from six manufacturers including Acer, ASUS and Dell; two units linked over ConnectX-7 pool into 128GB and support up to 200B parameters. For workloads where data can’t leave the premises (including the exfiltration risk in the previous item), running models locally is becoming a real option rather than a compromise.

Research radar

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

Controlled strong-to-weak distillation experiments on Llama 3 and Qwen 2.5 find that the rollout policy is not the main driver of performance: token-level KL direction shapes task performance and output coverage, while learning rate governs both catastrophic forgetting and update sparsity. The widely repeated claim that on-policy training is inherently better does not hold in their setup. Worth a close read if you pick post-training recipes by conventional wisdom.

A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review

The authors rewrote papers into 1,260 versions with identical science and different rhetoric, then tested 30 AI-reviewer configurations: scores move with the writing, and reviewers that seem insensitive to rewrites often get there by scoring every paper about the same, which the paper calls false robustness. Their SciCore reviewer reads both the full manuscript and an extracted scientific core, and does better on stability and discrimination at once. If an LLM assigns scores anywhere in your evaluation pipeline, this is a direct warning that it may be rewarding prose, not substance.

Today in one line: Ataraxos spent under $8,000 to reach a milestone that DeepMind’s 2022 attempt, estimated at $3M–$4.5M in today’s prices, never achieved; once frontier results fit a small lab’s budget, capabilities will spread faster than any governance framework built on the assumption that only big labs can afford them.