My day job is content moderation. Every decision I make starts with a taxonomy. A piece of content comes in; I determine which category it falls under (violence, harassment, misinformation); the category then determines the action: remove, reduce distribution, add a warning, or leave it up. Meta has described exactly this remove/reduce/inform structure since 2018 and in its current enforcement documentation. In this industry, a taxonomy is not an academic artifact. It is a standard operating procedure on a production line.

So I pay close attention to the state of AI risk taxonomies. As of May 2026, more than 70 of them had been publicly released (a count the paper I discuss below cites from Slattery et al.; its own appendix is explicitly a non-exhaustive list). MIT’s AI Risk Repository, which exists just to aggregate them, indexes 65 frameworks and over 1,700 individual risk entries. The taxonomies keep multiplying, and there is little evidence that any of them are wired into deployment decisions or post-launch monitoring. A paper posted to arXiv in August puts this on the table under the title “Death by a thousand taxonomies?”: what kills governance may not be the one risk that slips through a category, but the thousand incompatible classification schemes themselves.

What the paper did

Glen Berman and five co-authors interviewed 25 people who have actually built or used AI risk taxonomies: 8 from academia, 10 from industry, 7 from civil society and government. Seventeen of them had contributed to developing at least one taxonomy, more than 20 in total. Field checks like this are rare; as the authors note, there has been little empirical research on how these schemes actually get developed, interpreted, and used.

A definition first. The paper calls its object a Sociotechnical Outcome Taxonomy, or SOT. In plain terms: a structured classification of risks, impacts, and harms, typically organized as top-level categories such as privacy violations, misinformation, discrimination and exclusion, each split into finer subcategories, meant to be consulted during evaluations, audits, and policy writing. Regulators have their versions, many companies build their own, and researchers keep publishing new ones.

The interviews compress into three diagnoses.

First, the design trade-offs are invisible, so users mistake the list for the full set of risks. Every taxonomy gets carved out by specific people under specific constraints: who was in the room, how fine the categories run, whose workflow it serves. Those choices rarely appear in the finished product. One interviewee put it bluntly: “If the taxonomy is created primarily from the … circle jerk of computer science, then it’s not necessarily actually mapping the full breadth of risk.” Another offered an example: a deepfake voice used to fool a grandmother, versus an AI-cloned voice used to impersonate Secretary of State Marco Rubio in messages to at least three foreign ministers (that one actually happened, in June 2025; the State Department opened an investigation after it surfaced in July). The harm mechanisms differ entirely. One is fraud against an individual; the other is an impersonation operation aimed at diplomats. A taxonomy with a single deepfake-harm category presses both into the same cell. Users who never see the compression treat the cell as the reality. In interviewee I8’s words: “The danger, of course, is always that you start to mistake your abstraction for reality … people think the taxonomy is real.”

Second, they list harms without attaching responsibility. This is the finding I recognize most. I8’s line deserves quoting in full: “There’s a missing element of causation in the taxonomy … We are treating it like a botany taxonomy as opposed to, ‘Wait, we’re the ones doing this.’ … Whenever it’s discovered or detected and pointed out, of course they … go to court and say, ‘No, we’re not responsible.’” A botanist classifying plants bears no responsibility for the plants’ existence. The people who build and use AI risk taxonomies are the ones producing the risks, yet the taxonomy adopts a bystander’s posture: it records which harms exist, not which decision point and whose choice let them happen. That is where accountability washes out.

Third, classification is standing in for action. The scenes interviewees describe are blunt. A product team completes a datasheet (a document recording a dataset’s provenance and composition), gets asked what they changed because of it, and answers: “We didn’t actually change it. We just document.” Another interviewee: “I use this tool and this sort of Bible, in quotes, oh, I’m done … It’s okay. I got the green light. Push it into market.” Post-launch risk monitoring “consistently gets de-prioritised”; an adverse event reporting pipeline is “like a fantasy.” Meanwhile, the sheer number of taxonomies generates work of its own: teams now build crosswalks between schemes so that work done under taxonomy A still counts under taxonomy B. Translation labor replaces governance labor. The paper’s term is analysis paralysis: choosing which scheme to use consumes the capacity to act.

Why they multiply: classification is narrative power

The paper doesn’t claim a single cause for the seventy-plus schemes, but it names one incentive directly: industry researchers get rewarded for publishing new taxonomies. There is also a deeper pull. One interviewee named it: “At the end of the day, taxonomies are language creation … if language and categorisation is power, taxonomies hold power. Because if you create [a] taxonomy and you leave out categories intentionally, those categories won’t get measured.” Define the taxonomy and you define the boundary of what counts as risk. Another observed that the ways some companies frame risk are “sometimes suspiciously narrow.” With no coordination mechanism, companies and research groups keep shipping their own versions, and the developers interviewed had little way to learn what happens afterward. One admitted: “I unfortunately have no evidence whether people have found it useful.”

The contrast from content moderation

Hold a content moderation policy taxonomy next to these and you can see what’s missing. Moderation taxonomies are not more scientific. In my experience, annotators disagree, borderline content stays subjective, and different user communities draw the same category’s line in different places. Calibration sessions, double annotation, and appeals review push disagreement rates down; I have never seen them reach zero. But moderation taxonomies have one property AI risk taxonomies generally lack: in every system I’ve worked in, each label connects to an action. Categories map to enforcement actions, actions have thresholds, every policy has an owner, decisions feed quality metrics, wrong calls flow through an appeals channel. The taxonomy is the first stage of an enforcement pipeline. It doesn’t need to be “adopted,” because without it the line doesn’t run. AI risk taxonomies sit in the opposite position: build the list first, then hope someone volunteers to wire it into decisions, except that execution layer doesn’t exist. As the paper puts it, realizing the potential of SOT requires governance infrastructure that doesn’t yet exist.

What would actually help

Two of the paper’s design recommendations strike me as practical. Interoperability: shared structural conventions (distinguish risks from impacts from harms, distinguish observed from anticipated effects) plus published crosswalks between schemes, turning the translation work practitioners now do privately into part of the release itself. Traceability: for each risk, record the decision points it travels through, who owns each of those decisions, and what post-deployment signal would reveal it. The paper calls this lightweight causal mapping. I’d call it the minimum edit that turns a botany catalog back into an operating procedure.

On the governance side, the paper proposes a shared registry where taxonomies are deposited alongside their design documentation, and argues explicitly against standardizing now: premature standardization would freeze incumbents’ definitional advantages into the standard. What’s needed first is the material a future standard could grow from — usage evidence, cross-taxonomy mappings, monitoring data.

For everyone else, a more direct use. Next time a company or a report presents a risk taxonomy, test it with three questions. What action does each category trigger? When a harm occurs, who at which step is responsible? After deployment, what signal detects it? A taxonomy that answers none of these, however finely it slices, is paperwork. And in my experience, the job that kind of paperwork ends up doing is to prove, after something goes wrong, that we did classify it.

References