Last October, OpenAI published a post on how ChatGPT handles mental health crises (Strengthening ChatGPT’s responses in sensitive conversations): more than 170 mental health professionals involved, undesired responses down sharply. Coverage mostly tracked the improvement. At least in the discussions I saw, almost nobody mentioned another number in the same post: agreement between those clinical experts ran 71 to 77 percent. In roughly a quarter of cases, the experts disagreed about whether a response was acceptable.
OpenAI reported that number as background. A Stanford team made it the research question: if the experts don’t agree in the first place, what is a “safety evaluation” built from their scores actually measuring?
The paper (arXiv:2601.18061, accepted at ACM FAccT 2026, the ACM conference on fairness, accountability, and transparency) gives an uncomfortable answer: the disagreement is structural and can’t be treated as noise, and the “ground truth” you get by averaging it away belongs to no expert’s clinical judgment.
What the study did
The team wrote 90 synthetic prompts, with no real user data: ten high-risk categories (endorsing suicidal ideation, supporting non-suicidal self-injury, reinforcing hopelessness, validating hallucinations, playing along with delusions, encouraging manic behavior, encouraging avoidance, supporting substance use, reinforcing compulsive behavior, reinforcing distorted body image), plus a low-risk control category, penalizing ADHD-typical behavior. Each category spans nine variants: three levels of clinical severity crossed with three levels of directness. The severity levels are anchored to clinical instruments like the DSM-5 and the Columbia Suicide Severity Rating Scale (C-SSRS), not invented on the spot.
The 90 prompts went to four models: GPT-5, Claude 4 Sonnet, Grok-4, and Llama 3.2, for 360 responses. Three board-certified psychiatrists then scored each response independently on a calibrated eight-factor rubric: two safety factors (severity of potential harm, likelihood of harm) and six quality factors (clinical correctness, relevance, active listening, empathy without stigma, boundaries and disclaimers, actionability), on a five-point scale.
The standard tool for measuring whether raters agree is the intraclass correlation coefficient (ICC; 1.0 is perfect agreement, and the common convention treats anything under 0.5 as poor reliability; this paper used a more lenient 0.40 cutoff). Result: all eight factors landed between 0.087 and 0.295. A second reliability measure, Krippendorff’s α (0 is chance level; the commonly used floor for usable data is 0.667), looks worse: four factors came out negative, with boundaries and disclaimers at −0.203, less consistent than random scoring. Not one factor passed.
Two objections come naturally. Sample too small? The researchers ran the same questions as a live poll at the American Psychiatric Association’s 2026 annual meeting, with over 100 psychiatrists in the room; the split persisted (Stanford HAI’s writeup). One rater phoning it in? Random noise would scatter evenly; instead, disagreement concentrated in the highest-risk categories. Scores on suicide and self-harm responses were more dispersed (mean absolute deviation 0.598 and 0.566) than on the low-risk ADHD control (0.461). Exactly where expert consensus matters most is where it fails.
Where the disagreement comes from: orientation, not competence
The paper has a concrete case. On the boundaries-and-disclaimers factor for the same set of responses (did the model make clear it isn’t a therapist, should it have referred the user to professional help), one psychiatrist gave 92 percent of responses a score of 2, a near-blanket fail; another spread their scores between 3 and 5, mostly passing. Neither misread anything. They were fighting over a normative question: how explicitly must an AI state its limits to a user in emotional crisis? One camp holds that failing to state them, repeatedly, is negligence; the other holds that boilerplate disclaimers push away someone who has just worked up the nerve to talk.
The authors sort the disagreement into three mutually incompatible clinical orientations: safety-first (refer out and state limits even at the cost of interrupting the conversation), engagement-centered (preserve the relationship first, so the person keeps talking), and culturally informed (the same sentence carries different risk signals, and calls for different responses, in different cultural contexts). The same response gets opposite scores: the safety-first psychiatrist faults it for not referring out, the engagement-centered one credits it for not driving the person away. As co-author Nina Vasan, a clinical assistant professor of psychiatry at Stanford, put it: “The disagreement is structural, not just noise or even bias in the data.” “It’s a matter of professional judgment.”
Why the average is harmful
First layer: an average is not a consensus. The paper’s framing is that when disagreement is this systematic, averaging manufactures a new value framework that no individual expert holds. Two physicians, one prescribing drug A and one prescribing drug B, don’t add up to half a dose of each; that’s a prescription neither would write.
Second layer: these labels get inherited downstream. RLHF (training models on human preference scores) depends on reward models; train one on averaged labels like these and it learns a blend of individual rater habits and actual response quality. The same goes for safety classifiers in the Llama Guard family, which are trained on human annotations and deployed in products to block dangerous content. The paper’s point is that these training labels are least reliable precisely in the domains where safety classification matters most.
Third layer, and to me the most practical: it drains meaning out of comparisons between safety scores. The aggregate ICC was 0.269. Extrapolating from that (and it is an extrapolation: the study used the same panel of three throughout, and the authors themselves call for replication with larger, differently composed panels), a different but equally qualified panel of psychiatrists could hand the same model a visibly different score. When a vendor report says model A beats model B on safety by a few points, that gap may be substantially a fact about who did the scoring. Without inter-rater reliability disclosed, the score can’t be interpreted.
Content moderation learned long ago that consistency is manufactured
My day job is content moderation, and this problem is my daily scenery. The moderation industry never assumes annotators naturally agree. Consistency is built by institutions: first a policy document, often running to dozens or hundreds of pages, pins the boundary down; then calibration sessions, QA sampling, and escalation paths grind people toward alignment. Even then, at least in the work that has crossed my desk, disagreement on boundary cases never really disappeared. A written rule adjudicates it by fiat. And where that rule gets drawn, how false positives trade against false negatives, is not something measurement can answer. It’s a product and values decision, and the answer differs across user populations and cultural contexts.
Held against that, the “have a few experts score it and average” style of AI safety evaluation quietly skips the hardest step in the moderation pipeline. Someone is supposed to decide which standard applies and be accountable for it. Instead, the arithmetic mean makes the value decision, and leaves no record that a decision was made.
The paper’s recommendations boil down to three, and all of them put that step back. Preserve the disagreement instead of averaging it out; as first author Kiana Jafari, a Stanford postdoc, put it, “when they do not agree, you are not actually getting to the ground truth by averaging their scores.” Require developers to disclose which clinical framework an evaluation used. Model the frameworks separately rather than squashing them into one number. To be clear, none of this makes the disagreement go away. The three clinical orientations are incompatible and can’t be meaningfully averaged, and a global product faces cultural variation wider still. This is a difficulty in the problem itself; better statistics won’t dissolve it. What these practices do is put the choice on the table, so that “whose standard did you adopt” becomes an explicit decision someone can be asked about.
The study’s own boundaries need stating too. All 360 responses answer a single prompt; what got evaluated is single-turn replies. Real crisis conversations run over many turns, and this study didn’t test that. Multi-turn safety is being worked on elsewhere, for instance The Slow Drift of Support, which measures support boundaries loosening across turns of mental health dialogue, and the clinician-annotated multi-turn crisis dataset CRADLE-Dialogue. But whether expert disagreement widens or converges in multi-turn settings, I couldn’t find a study that measures directly; that’s an open gap. Three raters is genuinely few. The APA poll of 100-plus psychiatrists isn’t a rigorous replication, but it split the same way, so “add more experts and it converges” did not happen on the one occasion it was tried.
Conclusion
In my piece on the MIT Sloan study of AI financial advice (the control group didn’t contain a human advisor), the point was that evaluation design decides how far a conclusion can be read. This study goes a level deeper: the ground-truth layer itself may not exist. Next time you see a safety claim of “reviewed by N clinical experts,” ask three things: how many experts, what was their agreement rate, and which clinical framework. OpenAI at least printed its 71–77 percent; the paper recommends that evaluations of this kind report inter-rater reliability as a rule, and I’d treat that as the disclosure floor for the industry. Behind a safety score there is no neutral measurement, only a position that is either declared or hidden.
References
- Stanford Study Exposes Major Flaw in AI Mental Health Safety Testing | Stanford HAI — story origin; APA annual meeting live poll of 100+ psychiatrists, the three clinical orientations, Jafari (first author, Stanford postdoc) and Vasan (clinical assistant professor of psychiatry) roles and quotes, “three board-certified psychiatrists rated 360 responses,” no real user data, FAccT 2026 acceptance
- Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing (arXiv:2601.18061) — the paper; four models (GPT-5, Claude 4 Sonnet, Grok-4, Llama 3.2) × 90 prompts, high-risk category list and 3×3 design, eight-factor rubric, ICC 0.087–0.295 (aggregate 0.269, cutoff 0.40), four factors with negative α (boundaries factor −0.203, floor 0.667), suicide/self-harm mean absolute deviation 0.598/0.566 vs ADHD 0.461, the “92 percent scored 2” case, the inference about RLHF reward models and Llama Guard-style classifiers, the point that averaging produces a value framework held by no one, the recommendation to report inter-rater reliability
- Strengthening ChatGPT’s responses in sensitive conversations | OpenAI — 170+ mental health professionals involved, 71–77 percent inter-rater agreement among clinical experts (page blocks scraping and could not be opened directly; figures cross-checked against multiple secondary reports, see fact-check notes)
- OpenAI strengthens ChatGPT mental health guardrails | Becker’s Behavioral Health — secondary confirmation of OpenAI’s 71–77 percent agreement figure
- A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research (Koo & Li 2016) — ICC interpretation conventions (below 0.5 is poor reliability)
- The Slow Drift of Support (arXiv:2601.14269), CRADLE-Dialogue (arXiv:2606.10380) — adjacent work on multi-turn mental health dialogue safety, cited to note that no study directly measures how expert disagreement changes with turn count