OpenAI reaffirms zero data retention, previews “Private Safety Processing”

OpenAI said on August 19 it will keep offering Zero Data Retention to eligible API customers: once a request is processed, neither prompts nor outputs are stored, with one legally mandated exception: content flagged as suspected CSAM is retained for review and reporting. The new part is a preview of Private Safety Processing, rolling out in September alongside a technical white paper. Some recent frontier-model deployments have required retaining customer content so staff can inspect it when abuse is suspected; OpenAI says the new system lets automated classifiers look for risk patterns across related interactions and send back only a narrowly defined safety signal, with no employee access to customer content. Whether this works comes down to how that signal is defined: narrow enough to count as private, broad enough to catch abuse. Until the white paper answers that, enterprises have nothing to verify.

Vetted security researchers briefly lost access to OpenAI’s trusted cyber program

Trusted Access for Cyber gives identity-verified defensive researchers access to models with relaxed cybersecurity guardrails; a new tier called Daybreak Blue, with frontier-model access, launched August 10. On August 19, multiple researchers found their access gone, with messages saying their identity could not be verified or their accounts were ineligible. The at least five researchers who spoke to TechCrunch all live outside the US and Europe. OpenAI called it a technical error on its end and asked them to re-verify their identities. My read: OpenAI is pacing model development around cyber capabilities and steering defenders toward this official channel, which makes the access system itself part of the security infrastructure. If defenders’ workflows depend on tiered access, outages need response plans and recovery timelines, the same way API uptime does.

Adoption keeps climbing, approval doesn’t: the US polling gap on AI

TechCrunch pulled together recent polls: Pew found 52% of Americans are more concerned than excited about AI in daily life, up from 37% in 2021; an Economist/YouGov survey in May had 71% saying AI is advancing too quickly; a CNBC poll found most 18-to-34-year-olds don’t trust AI industry leaders to act responsibly. The piece argues this is a real trust deficit rather than a messaging problem: people use AI constantly but experience it as job risk and unwanted features, with the benefits landing somewhere other than their own lives. Airbnb CEO Brian Chesky wants “more regular things” that prove value; Anthropic CEO Dario Amodei concedes “we haven’t yet delivered on our big promises to benefit the world. That is totally on us.” Penetration is pushed by vendors; approval is granted by users. My bet is that this gap eventually turns into regulation and churn.

Stanford HAI: open weights aren’t open source, and science needs the real thing

Stanford computer science professor James Landay argues that releasing weights alone is open distribution, not openness: you can run the model, but you can’t see how it was built, what data trained it, or why it behaves the way it does. His standard, drawn from the Linux Foundation’s Model Openness Framework, requires open weights, open training data (or fully auditable documentation of it), and open code and tooling, and he wants universities to build such models on longer timelines than commercial cycles allow. For research the argument holds: without the training data you can’t study how that data shaped the model’s behavior. Open training data, though, carries copyright and privacy exposure, and who absorbs that cost decides how far this goes. The university proposal is, in effect, a bid to move that cost into the public sector.

Stanford Law’s Daniel Ho, Columbia Law’s Neel Guha, and colleagues published “There’s No Free Benchmark” in PNAS, arguing legal AI lacks what they call legibility: lawyers, judges, and clients have almost no way to know these systems’ real error rates or where they fail. Documented court cases involving AI-fabricated facts, precedents, and statutes now exceed 1,700. The paper treats benchmarking as an institutional problem: what gets measured is shaped by the incentives and resources of developers, firms, academics, and regulators, and low-resource areas like bankruptcy and child custody, the kinds of legal problems where ordinary people are most likely to turn to a general chatbot for help, get little benchmark coverage. The authors propose institution-specific fixes, including a role for independent bodies like NIST. The diagnosis, that evaluation gaps come from incentive structures, applies well beyond legal AI.

Bloomberg: SpaceX explored acquiring Cognition; the CEO says “not for sale”

Bloomberg reported that SpaceX tried to acquire Cognition, the company behind the coding agent Devin, and that talks are no longer active though the two still discuss cooperation, such as Cognition using SpaceX compute. CEO Scott Wu disputed the story on X within hours, saying it was inaccurate, the company is not for sale, and no talks happened (see also TechCrunch). The backdrop: SpaceX closed its $60 billion acquisition of Cursor’s parent Anysphere last week. Whatever actually happened, coding-agent companies are now assets that non-software giants shop for, and for a startup with Mercedes-Benz, Citi, and Goldman Sachs as customers, saying “not for sale” out loud is itself a stability signal to clients and employees.

Research radar

HarnessRisk: a lifecycle benchmark for agent harness safety

An agent harness is the layer that manages an agent’s tools, extensions, persistent state, permissions, and external actions. This benchmark pairs 128 legitimate user tasks with adversarial instructions hidden in untrusted workflow components, across six lifecycle phases from configuration through incident recovery. Attack success rates ranged from 12.6% to 80.9% across systems, configuration was the weakest phase, and the finding worth remembering is that detection does not imply safety: some configurations flagged the risk in over 90% of runs and let the attack through anyway. Read it if you build or secure agent harnesses.

Demystifying agent skills: why they work, until they don’t

Controlled experiments over 8,135 trial records on agent skills, the structured knowledge packages you attach to an agent. Skills mostly work through procedural anchoring, stabilizing the agent’s noisy execution on known steps (65.7% of skill cases); injecting missing knowledge accounts for only 4.5%. Retrieval collapses as the pool grows: 29.6% precision with 5 skills, 3.3% with 100. The practical translation: write skills as procedures instead of knowledge docs, and keep the library small.

Agentic ESOpt: fine-tuning long-horizon agents on inference-level GPU memory

Replaces RL with evolution strategies for full-parameter fine-tuning of long-horizon agents: sample parameter perturbations, evaluate each variant on the task, apply reward-weighted updates, no backpropagation. That cuts GPU memory to inference level and sidesteps credit assignment under sparse long-trajectory rewards, since attribution happens at the whole-trajectory level. On WebArena-Lite it improves Qwen-3.5-27B by 6.69%. Relevant if you train agents without a large GPU budget.

Fool’s Gold: if you can’t stop safety removal, poison what it yields

Abliteration, which finds and projects out a model’s internal “refusal direction,” strips an open-weight model’s safety alignment in minutes and is essentially unpreventable. This paper (by Mark Russinovich) stops trying to prevent it and instead trains “decoy hardening” inside a differentiable simulation of the attack: once alignment is stripped, the model answers dangerous requests fluently and confidently but with the critical details falsified; unattacked behavior is unchanged. Six of seven models tested (9B to 122B) passed, with 51% to 90% of attacked-state responses being decoys. The stated limits: repeated sampling partially recovers usable answers, and in-context jailbreaks are out of scope. It reframes open-weight safety from preventing tampering to making tampered outputs untrustworthy; worth a close read for open-model safety and red-team researchers.

Today in one line: Recent frontier deployments have treated zero data retention and safety monitoring as an either/or; OpenAI now claims both, and until September’s white paper shows a checkable mechanism, that claim is a product promise, not an established fact.