On August 19, OpenAI published an announcement: Zero Data Retention (ZDR) will continue to be offered for frontier models, alongside a preview of a new mechanism called Private Safety Processing. Read in isolation, it looks like a routine enterprise feature update. Read in context, it is a direct engagement: Anthropic announced that existing zero-retention agreements no longer cover its Mythos-class models (Claude Mythos 5 and Claude Fable 5). Enterprises that want those models must accept 30-day retention of prompts and outputs, under a policy that took effect June 9. TechCrunch’s headline put it plainly: “OpenAI seeks to one-up Anthropic.”
Same question, opposite answers from the two leading labs: does abuse monitoring for frontier models require keeping customer data? This post takes apart each lab’s technical route, and asks how “scan for harmful use” and “protect data privacy” can coexist as an engineering matter.
The announcement details below are cross-checked against OpenAI’s official X post and reporting from Axios, TechCrunch, and Computerworld.
What zero retention promises, and what it doesn’t
ZDR is a contractual arrangement between an API provider and enterprise customers it has approved: on qualifying endpoints, the prompts you send and the outputs the model generates are not retained once inference completes. OpenAI’s data-usage documentation defines this boundary, including its explicit carve-outs (suspected CSAM among them) and the prior approval it takes to get it. Note what it does not cover. In the standard hosted-API setup, your request still travels to the provider’s servers and is processed there in readable form. ZDR governs what happens to the data after processing, not the processing itself.
For customers who carry their own confidentiality obligations, law firms, hospitals, financial institutions, retention terms are a natural sticking point in compliance review. Whether 30 days of data sitting with a third party is acceptable depends on the specific contract, the jurisdiction, and the data type; there is no general answer. Which is exactly why “we keep nothing” is the easiest position to defend in that review, even though, as we will see, zero is never quite literal.
Where the conflict lives: abuse monitoring wants history
The provider’s abuse monitoring (per OpenAI’s data-usage documentation: automated classifiers plus logs that can include prompts, responses, and classifier outputs, kept to enforce Usage Policies and mitigate harmful uses) has a need that collides head-on with zero retention: it wants to look back.
A classifier operating at the level of a single message sees only isolated requests. TechCrunch describes the general pattern as bad actors who “spread out their requests to avoid detection.” A hypothetical (this is my construction, not from the announcement): split one bad project into pieces. Today, in one session, ask for network-scanning code. Tomorrow, in a fresh session, ask about privilege escalation. The day after, persistence mechanisms. Each request reads as an ordinary programming question; assembled, they are an intrusion-tool development pipeline. Detecting that pattern requires looking across sessions. And the more agentic AI becomes, with longer-running tasks, more autonomy, and more interactions per task, the sharper this need gets. OpenAI’s own X post makes the argument in exactly these terms: “As AI takes on longer, more autonomous work and delivers greater value to businesses, safety systems also need to identify risks across related interactions.” Anthropic’s stated rationale for its 30-day retention is likewise safety work.
So the conflict sits in the open: customers want nothing kept, safety teams want data they can study. The two labs differ on who yields to whom.
Anthropic’s route: keep the data, govern who sees it
Anthropic’s answer is governance. Per the official policy page, inputs and outputs for Mythos-class models are retained for 30 days, then deleted automatically (except content flagged by safety systems or subject to legal holds). By default, staff cannot access the data. Human review “can occur only through a controlled access path,” triggered when automated trust-and-safety systems flag potentially harmful content. Every access is recorded in a tamper-proof log that reviewers cannot suppress or modify. Eligible organizations can add customer-managed encryption keys and access transparency audit logs. The requirement covers every channel, including zero-retention customers on AWS Bedrock, Google Cloud, and Microsoft Foundry; decline the retention, and the models are unavailable.
Note what customer-managed keys actually do in this design. The data still does not sit with the customer: it lives on Anthropic’s side for direct API use, or stays within the cloud platform’s environment for Bedrock and Google Cloud channels. Once content is flagged, a reviewer can still read the flagged conversation through the controlled path. The whole scheme amounts to “the data sits with the vendor, and process guarantees that only the right people see it at the right time.” What you trust is the process.
OpenAI’s route: move the plaintext and the keys
Private Safety Processing takes a different path: it does not touch “whether to keep” and instead changes where the history lives and who holds the keys. As Axios describes it, customers get two deployment options. One: the historical data used for safety analysis stays on infrastructure the customer controls. Two: OpenAI stores it, encrypted with customer-controlled keys, with no key copy on OpenAI’s side and no way for staff to decrypt. Automated systems run cross-interaction pattern analysis over this data, and when they detect suspected misuse, they send OpenAI only “a narrowly defined signal” warning of a specific type of activity, without the prompts and responses themselves. Per TechCrunch, OpenAI then decides whether enforcement is needed, and the customer “may choose to share data with OpenAI at their discretion” to help resolve the issue.
Here is the key technical question the announcement leaves open. To judge whether content is dangerous, a classifier has to work on the content itself, decrypted into some computable form somewhere. If that content sits on customer infrastructure, or in storage locked with customer keys, then where does the classifier run, and who decrypts for it? I can think of three implementations (speculation on my part; the announcement doesn’t say): the classifier ships as a component running inside the customer’s infrastructure; it runs in a confidential computing environment (a TEE, or trusted execution environment: a processor-enforced encrypted enclave where decryption happens only inside the enclave, and even the machine’s operator and the cloud platform cannot read enclave memory); or scoring happens synchronously at inference time, persisting only scores and never content. The three carry entirely different trust models. OpenAI says the rollout begins in September along with a technical white paper. Until then, I can verify what is promised, not what is built.
Two openings show that “zero” was never literal. First, the legal exception: even in zero-retention deployments, images flagged as suspected child sexual abuse material (CSAM) are still retained for human review and legally required reporting. Brian Levine, a consultant and executive director of FormerGov whom Computerworld interviewed, is blunt about it: “Zero is never quite zero.” Second, the safety signal itself is information distilled from customer data and handed to OpenAI. How narrowly the signal’s fields are defined determines how much this channel can carry out. That is the second question I want the white paper to answer.
The shape of this idea is familiar: content stays on the user’s side, and only a verdict on whether it matched a dangerous pattern travels back to the vendor. Apple’s client-side CSAM scanning plan, proposed in 2021 and abandoned by the end of 2022, pointed the same direction. The two are not technically equivalent: Apple’s design was on-device hash matching with account-level thresholds (30 matching images before Apple’s server-side review kicked in), not cross-interaction behavioral classification, so the analogy only goes so far. But the objections from back then apply unchanged today. The joint paper from security researchers centered on one point: once a scanning channel exists, it can be extended into surveillance beyond its originally stated purpose, and the matching rules can quietly widen. Enterprises evaluating Private Safety Processing should ask the same question.
How to use this news
The two labs look opposed on the surface, but their directions are converging: in both designs, content goes to automated systems first, and the human role shrinks. It does not reach zero in Anthropic’s design: a small set of reviewers can still read flagged conversations through the controlled path. OpenAI pushes further: per the announcement, staff cannot read even flagged content; they see only the signal. The remaining disagreement narrows to one variable: whose hands hold the cross-session plaintext history, and whose hands hold the keys. Anthropic’s answer is “it sits with us, and we govern the process.” OpenAI’s is “it sits with you, or it sits locked and you hold the key.”
For enterprise customers, this changes how compliance evaluation should be phrased. The old review examined the vendor’s policy documents. The new review should ask three technical questions in order: where is the plaintext history physically stored; who can decrypt it, and in what execution environment does decryption happen; and what fields, exactly, does the signal that leaves the customer boundary contain. Under Anthropic’s model, the answers live mostly in policy language and audit clauses. Under OpenAI’s model, if the white paper delivers what is promised, part of the answers can live in architecture, verifiable by third parties. Which layer the verifiability lands on is the essential difference between the two routes.
But it is too early to score this today. What OpenAI has shipped so far is a promise and a preview. The eligibility criteria for “qualifying enterprise and API customers” are unpublished: OpenAI’s own documentation says only that the controls require prior approval and points prospects to its sales team, and Computerworld flagged the gap specifically. The classifier’s execution environment, the signal’s exact definition, and the handling of false positives are likewise unanswered in public; September’s white paper is the earliest chance for answers, and there is no guarantee it covers all of them. If the paper is vague, this announcement was a marketing move aimed at Anthropic. If it is detailed enough to audit, it puts pressure on abuse monitoring across the industry: once “no human can see the content” is demonstrably buildable, the default equation of “safety, therefore we must keep your data” stops being a premise anyone can treat as self-evident.
My read is that the second possibility deserves to be taken seriously. See you in September, white paper in hand.
References
- Offering Zero Data Retention for frontier models — OpenAI — the announcement (details cross-checked against the official X post and the three reports below)
- OpenAI’s official X post — verbatim source for “We will continue to offer Zero Data Retention for frontier models” and safety systems needing to “identify risks across related interactions”
- Your data — OpenAI API documentation — official definitions of ZDR (eligible endpoints, prior approval, carve-outs) and abuse-monitoring logs
- Data retention practices for Covered Models — Anthropic Privacy Center — full detail on the 30-day policy: covered models, effective date (June 9, 2026), staff access restrictions, tamper-proof logs, customer-managed key option
- OpenAI previews zero-retention safety system as Anthropic requires data logs — Axios — the two deployment options (customer infrastructure / customer-key encryption), the safety signal, September white paper, CSAM exception
- OpenAI seeks to one-up Anthropic with new customer privacy protections — TechCrunch — competitive framing, enforcement flow after a signal fires, voluntary data sharing by customers
- Computerworld report — unpublished eligibility criteria; consultant Brian Levine’s “zero is never quite zero” comment
- Apple confirms it has stopped plans to roll out CSAM detection system — 9to5Mac — timeline of Apple’s client-side scanning plan (proposed 2021, abandoned December 2022) and its on-device matching design
- Apple’s CSAM system details — 9to5Mac (2021) — source for the 30-image account-level threshold before server-side review
- Bugs in Our Pockets: The Risks of Client-Side Scanning — arXiv — analysis of how client-side scanning can be abused or extended into a surveillance tool