Google releases HEIR: cloud AI inference on data that stays encrypted
Google’s security team published HEIR, an open-source compiler toolchain that converts trained AI models into versions that run on homomorphically encrypted data. Homomorphic encryption means the math happens directly on ciphertext: the server never sees plaintext at any point, and the result it returns is also encrypted, readable only by the user who holds the key. Google shipped four demos alongside it: private content recommendation (with Belfort Labs, LG, and NYU), credit card fraud detection, network threat detection, and audio hotword spotting, plus partnerships with four hardware accelerator companies (Belfort, Niobium, Cornami, Optalysys). My read: a compiler is the right wedge, because it separates the people who understand cryptography from the people who understand models, and that separation is what moves homomorphic encryption from papers into engineering pipelines. But the latency numbers the post does show are for a single-threaded CPU; accelerator results are only promised for the near future, which tells you everyday products are still a good distance away.
GLM-5.3 ships: base model untouched, all gains from post-training
Zhipu (Z.ai) released GLM-5.3. The official story: the base model was left alone, and every capability gain came from scaling up post-training. Internal evals put coding 50% ahead of GLM-5.2, open weights are promised in two weeks, and the model is available now through the GLM Coding Plan, working with coding agents such as Claude Code. The release page’s title makes cyber capability the second selling point after coding. Nathan Lambert’s analysis is worth reading: he estimates GLM-5.3 at roughly 750B parameters, a third the size of Kimi K3, yet ahead of it on most benchmarks, which he argues rules out the “distilling American models” explanation. In his account, Chinese labs keep pace through release cadence (days, where US companies spend months on internal testing) and trade-offs like text-only specialization. A model that reaches the frontier by spending only on post-training is a concrete counterexample to the story that pretraining scale decides everything.
Google turns the visible AI watermark into a toggle
Google announced that over the next few days, users can switch off the visible watermark on AI-generated images (Nano Banana), video (Omni), and music (Lyria) in Gemini and its video tool Flow, except where the law requires it, with Search support coming later (announced only on X, no blog post; see also TechCrunch’s report). What goes away is only the corner badge you can see. The invisible SynthID watermark and C2PA metadata stay embedded in everything: SynthID is Google’s hidden marker, statistical patterns woven into pixels and audio that human eyes can’t spot; checking for it means running the file through a detector, for instance by asking Gemini whether an image carries the mark. C2PA is the industry metadata standard recording who made a file and with what tool. Gemini lead Josh Woodward framed it as “striking a balance between creative control and safety.” My view: this shifts the job of identifying AI content from anyone who glances at it to whoever takes the extra step of running a detection check, and most readers scrolling past an image never will. Anthropic’s official Claude text watermark page is also live (we covered the announcement earlier); both companies are converging on “invisible but verifiable.”
New forecast: the natural gas that data centers bet on could triple in price
Energy research firm Noreva forecasts that natural gas prices could triple in parts of the US over the next few years, with some trading hubs in Texas and Louisiana going past $10 per million BTU, against a current Henry Hub benchmark near $3. The mechanism: LNG export terminals keep expanding, connecting a domestic gas market that used to be relatively closed, and therefore cheap, to global prices. Noreva CEO Peter Gardett puts it plainly: “finally we’re connecting the domestic gas market to the global gas market.” Fuel is about half the cost of electricity from a gas plant, so a doubling or tripling feeds straight into AI data center power costs. Hyperscalers picked gas because it builds fast; fuel price exposure is the hidden bill for that choice. TechCrunch also cites survey data that 80% of US consumers already worry data centers are raising their utility bills, and a price spike would amplify that political pressure.
French startup Kog digs into GPU internals for another round of inference gains
Kog, an 11-person French startup, raised a seed round (co-led by Varsity VC, backed by Bpifrance and French Tech 2030). Its approach is to spend weeks or months of low-level engineering per GPU model, and it claims a path to 30x faster LLM inference on standard datacenter GPUs like AMD’s MI300X and Nvidia’s H200, aimed at workloads such as coding agents where speed is the product experience. To be clear about the evidence: the public demo so far is 3,000 tokens per second on a 2B-parameter model, the 30x figure is the vendor’s own claim with no independent verification, and the company itself set “10x on major models by September” as its next milestone. Agents have turned inference speed from a cost line into a user-facing property, so the demand is real; the conclusion waits on numbers from large models.
Research radar
How far can rhetoric bribe an AI reviewer?
The authors took 120 real anonymized ICLR 2026 submissions and built 4,200 controlled variants that change wording but not substance, then had five LLM reviewers score them. The effect is concentrated, not uniform: evidence framing and novelty stance move scores the most, and the direction depends on the reviewer’s baseline (low scores get pulled up, high scores pulled down, mid-range scores swing hardest). A stricter review protocol lowers average scores by 1.36 points but does not close the rhetorical hole. Conference organizers weighing LLM-assisted review, and anyone studying reward hacking, should click through: this is controlled evidence that AI judges can be attacked with writing style.
DarwinX: evolve the harness, not the model
DarwinX moves the target of agent self-improvement from model weights to the harness (prompts, tools, skills, control flow). It maintains a population of harnesses under a retain-and-extend rule: a variant is accepted only if it extends capability coverage without regressing other tasks, with weights frozen throughout. Against single-path greedy self-improvement, where a local win quietly breaks something else, the population-plus-archive scheme holds up better: Terminal-Bench 2.1 rises 7.7 points to 83.2%, WebArena-Infinity pass rate jumps from 43.5% to 93.0%, and evolved harnesses transfer to SWE-bench Verified unmodified. If you build agent systems, read it closely; nothing here touches weights, so the recipe drops straight into a production harness iteration loop.
Intern-S2-Preview: an agentic foundation model for science
A Shanghai AI Lab-affiliated team (125 authors) released a scientific agentic foundation model at up to 397B parameters, trained in three stages: multimodal pretraining on scientific corpora, a unified post-training pipeline (SFT, multi-task RL, agentic RL), and a set of inference optimizations. The goal is cross-modal reasoning over scientific evidence, interaction with scientific tools, and sustained progress on long-horizon tasks, with competitive results across science and agent benchmarks. Researchers working on research automation, especially in biology or cross-modal settings, can treat this as a substantive reference point for the AI-for-science path rather than one more general benchmark entry.
Today in one line: what you can see and what you can trust are splitting apart — computation runs on encrypted data, provenance hides inside pixels, and in both cases the judgment moves from human eyes to vendor-run detection tools.